Reliability over time of EEG-based mental workload evaluation during Air Traffic Management (ATM) tasks

Arico, Pietro; Borghini, Gianluca; Di Flumeri, Gianluca; Colosimo, Alfredo; Graziani, Ilenia; Imbert, Jean-Paul; Granger, Geraud; Benhacene, Railene; Terenzi, Michela; Pozzi, Simone; Babiloni, Fabio · 2015 · Crossref

DOI: 10.1109/embc.2015.7320063

archive: archived pipeline: cataloged verified

Get this paper ↗ (DOI — opens at the source; we link to it, we don't host it)

Summary

This study addresses the challenge of maintaining high reliability in EEG-based mental workload (MW) estimation over time, a critical issue for real-world applications in safety-critical environments like Air Traffic Management (ATM). While machine-learning techniques can assess MW with high temporal resolution, their accuracy often degrades across days due to day-to-day fluctuations in brain signals, necessitating frequent recalibration. The authors hypothesized that limiting the classifier to a small subset of neurophysiological features strictly correlated with MW—specifically frontal and occipital theta rhythms and parietal alpha rhythms—would prevent overfitting and ensure stable discrimination accuracy (DA) over time without daily recalibration. The experimental design involved twelve Air Traffic Control (ATC) trainees performing the LABY simulated ATM task under three difficulty levels: Easy, Medium, and Hard. Data were collected across three sessions: two consecutive days (Day 1 and Day 2) and one week later (Day 9). EEG signals were recorded using a 13-channel system, filtered, and segmented into 2-second epochs. Power Spectral Density (PSD) was calculated for the theta and alpha bands. A Stepwise Linear Discriminant Analysis (SWLDA) classifier was employed to select features. The researchers tested three feature selection constraints: using 5%, 50%, or 100% of the available spectral features. Cross-validation strategies included intra-day, short-term (Day 1 vs. Day 2), and medium-term (Day 1/2 vs. Day 9) comparisons to evaluate stability. The results demonstrated that the classifier’s performance depended significantly on the number of features used. When trained with 50% or 100% of the available features, the Area Under the Curve (AUC) values decreased significantly in both short-term and medium-term cross-validations, indicating a loss of reliability over time. In contrast, the classifier using only 5% of the features maintained stable AUC values across all cross-validation types, with no significant degradation in performance from Day 1 to Day 9. Furthermore, the EEG-based workload index (WEEG) derived from the 5% feature set successfully distinguished between the three difficulty levels (Easy, Medium, Hard) across all subjects and days, with significant differences observed between conditions. The study concludes that restricting the classifier to a minimal set of MW-related EEG features ensures stable mental workload estimation over a period of one week, eliminating the need for daily recalibration. This finding suggests that specific neurophysiological markers remain consistent for trained operators performing familiar tasks. However, the authors note limitations, including the short testing period and controlled experimental conditions. Future research is needed to validate this approach in real-world ATM scenarios over longer durations and to investigate whether these stable features generalize across different subjects.

Provenance

The full processing record for this entry. Every stage of this paper's journey through the pipeline is logged — what ran, with which tool and model, how many attempts it took, and when it last completed.

StageOutcomeToolModelPromptAttemptsCompleted
discover success Crossref 1 2026-08-09
archive success unpaywall 2 2026-08-09
extract success cached 3 2026-08-10
clean success clean 1 2026-08-09
chunk success chunk 1 2026-08-09
embed success embed Qwen/Qwen3-Embedding-8B 1 2026-08-09
enrich failed 1 2026-08-09
promote success 1 2026-08-09
summarize success llm qwen3.6-27b-nvidia summ-v5 2 2026-08-10
tag success vector_similarity 10 2026-08-11
verify success 2 2026-08-10

Summary generated by qwen3.6-27b-nvidia on 2026-08-10; verification: verified.

Topics

Ranked by relevance to this paper. Hover a topic for its definition.

Information type

What kind of knowledge this paper contributes, grouped by family — independent of topic (what it is about) and method (how it was studied).