Adaptive multimodal learning for driver cognitive state monitoring using transformer-based fusion with personalized meta-learning and federated optimization

Abinaya, G.; Dinakaran, K. · 2026 · Crossref

DOI: 10.1038/s41598-026-51635-3

archive: archived pipeline: cataloged

Get this paper ↗ (DOI — opens at the source; we link to it, we don't host it)

Summary

**Research Problem and Motivation** Driver fatigue and cognitive overload contribute significantly to global road fatalities, with drowsy driving accounting for a substantial portion of accidents in regions like India and the United States. Existing driver monitoring systems face limitations: camera-based methods fail in low light, vehicle-based metrics are delayed, and physiological sensors suffer from intrusiveness and inter-subject variability. Furthermore, centralized training of deep learning models raises privacy concerns under regulations like GDPR. This paper addresses these gaps by proposing an Adaptive Multimodal Learning (AML) framework for real-time cognitive workload assessment and fatigue detection that balances accuracy, personalization, and privacy. **Methodology** The study utilizes the CL-Drive dataset, which contains synchronized multimodal data from 21 participants across nine simulated driving scenarios of escalating complexity. The data includes EEG (4-channel, 256 Hz), ECG (3-lead, 512 Hz), EDA (128 Hz), and gaze tracking (50 Hz). The AML framework employs a hybrid 1D-CNN–BiLSTM architecture to extract spatiotemporal features from raw signals, with kernel sizes and filter counts tailored to the specific characteristics of each modality (e.g., capturing QRS complexes in ECG or alpha waves in EEG). These features are fused using a transformer-based network with cross-modal attention, which models interactions such as correlating gaze fixation losses with EEG theta-band surges. To handle individual variability and privacy, the framework integrates personalized meta-learning (MAML), allowing adaptation to new drivers with as few as five windowed samples (~10 seconds of data), and federated optimization with adaptive gradient compression to enable decentralized training without sharing raw biometric data. **Results** Experiments under strict cross-subject evaluation demonstrate the framework's efficacy. Under subject-independent 5-fold cross-validation, AML achieves 80.5% ± 1.8% accuracy on binary cognitive load classification without personalization, improving to 91.8% ± 1.2% with 20 calibration samples. Under the more rigorous leave-one-subject-out (LOSO) protocol, accuracy reaches 77.8% ± 2.6% without personalization and 84.0% ± 1.8% with 20 samples, outperforming the strongest published LOSO baseline by 1.6 percentage points. The transformer-based fusion yields a 3.6 percentage-point absolute accuracy improvement over the strongest conventional fusion baseline. The federated component reduces per-client data transfer by 38% through adaptive gradient compression. Additionally, the model exhibits robustness to sensor noise, maintaining 81.5% ± 2.3% LOSO accuracy with only five calibration samples per new driver. **Significance** This work advances intelligent vehicle safety systems by providing a privacy-aware, scalable blueprint for adaptive multimodal learning in human-centric AI. By enabling rapid adaptation to new drivers with minimal calibration data and preserving data privacy through federated learning, the framework addresses critical barriers to real-world deployment in in-vehicle monitoring systems. It offers a robust solution for detecting cognitive states in diverse populations, potentially reducing accident rates caused by fatigue and overload.

Provenance

The full processing record for this entry. Every stage of this paper's journey through the pipeline is logged — what ran, with which tool and model, how many attempts it took, and when it last completed.

StageOutcomeToolModelPromptAttemptsCompleted
discover success Crossref 1 2026-08-09
archive success canonical_url 1 2026-08-09
extract success cached 4 2026-08-23
clean success clean 1 2026-08-09
chunk success chunk 1 2026-08-09
embed success embed Qwen/Qwen3-Embedding-8B 1 2026-08-09
promote success 1 2026-08-09
summarize success llm qwen3.8-27b-gittensor summ-v5 3 2026-08-23
tag success vector_similarity 11 2026-08-11
verify success 2 2026-08-09

Summary generated by qwen3.8-27b-gittensor on 2026-08-23; verification: pending re-verification.

Topics

Ranked by relevance to this paper. Hover a topic for its definition.

Information type

What kind of knowledge this paper contributes, grouped by family — independent of topic (what it is about) and method (how it was studied).