Machine learning performance in EEG-based mental workload classification across task types: a systematic review
DOI: 10.3389/fnrgo.2025.1621309
archive: archived pipeline: cataloged
Get this paper ↗ (DOI — opens at the source; we link to it, we don't host it)
Summary
This systematic review addresses the lack of standardized benchmarks in electroencephalogram (EEG)-based mental workload (MWL) classification, which hinders the comparison of machine learning (ML) model performance across studies. The authors conducted a PRISMA-guided literature search across Web of Science, IEEE Xplore, PubMed, and Scopus, identifying 83 peer-reviewed studies published primarily between 2017 and 2024. The review categorizes these studies by task type (single-tasking vs. multitasking) and MWL labeling method (task load-based, subjective, or performance-based) to analyze how these factors influence ML classification accuracy. Methodologically, the review evaluates models ranging from classical algorithms (SVM, KNN, Random Forest) to deep learning architectures (CNN, RNN, Transformers). It assesses performance based on classification accuracy, the number of MWL levels, EEG segment length, and model robustness (random split, cross-session, cross-subject, or cross-task). A significant portion of the literature (37%) relied on random data splits, which the authors note is susceptible to data leakage and overestimates generalizability. Most studies utilized fewer than 2 seconds of EEG segments and classified only two MWL levels, limiting granularity. The primary finding is a significant drop in MWL classification accuracy in multitasking studies where MWL was rated based on quantitative task load, compared to single-tasking studies or those using subjective ratings. This suggests inherent challenges in estimating MWL in complex, real-world multitasking environments. Additionally, the review highlights that while EEG contains more MWL information than other physiological signals, spectral features vary by task type (e.g., frontal theta increases in memory tasks vs. alpha desynchronization in visual tasks), making universal metrics difficult. The analysis also identifies methodological flaws, such as insufficient clarity in task difficulty adjustment and lack of cross-task robustness testing. The significance of this work lies in providing the first comprehensive examination of ML performance across task types, highlighting that current models struggle with the dynamic cognitive demands of multitasking. The review underscores the need for standardized experimental protocols and robustness testing to improve the practical applicability of EEG-based MWL estimation in real-world scenarios where tasks are rarely isolated. It identifies open questions regarding cross-task generalization and the optimal balance between temporal resolution and classification accuracy, guiding future research toward more robust and generalizable models.
Provenance
The full processing record for this entry. Every stage of this paper's journey through the pipeline is logged — what ran, with which tool and model, how many attempts it took, and when it last completed.
| Stage | Outcome | Tool | Model | Prompt | Attempts | Completed |
|---|---|---|---|---|---|---|
| discover | success | Crossref | — | — | 1 | 2026-08-09 |
| archive | success | canonical_url | — | — | 1 | 2026-08-09 |
| extract | success | cached | — | — | 5 | 2026-08-23 |
| clean | success | clean | — | — | 2 | 2026-08-10 |
| chunk | success | chunk | — | — | 2 | 2026-08-10 |
| embed | success | embed | Qwen/Qwen3-Embedding-8B | — | 2 | 2026-08-10 |
| promote | success | — | — | — | 1 | 2026-08-09 |
| summarize | success | llm | qwen3.8-27b-gittensor | summ-v5 | 3 | 2026-08-23 |
| tag | success | vector_similarity | — | — | 17 | 2026-08-11 |
| verify | success | — | — | — | 2 | 2026-08-09 |
Summary generated by qwen3.8-27b-gittensor on 2026-08-23; verification: pending re-verification.
Topics
Ranked by relevance to this paper. Hover a topic for its definition.
Information type
What kind of knowledge this paper contributes, grouped by family — independent of topic (what it is about) and method (how it was studied).
- Empirical Findings: physiological data