EEG-based Decoding of Auditory Attention to Conversations with Turn-taking Speakers
DOI: 10.1101/2025.06.20.660726
archive: archived pipeline: cataloged verified
Get this paper ↗ (DOI — opens at the source; we link to it, we don't host it)
Summary
This study addresses the limitations of traditional Auditory Attention Decoding (AAD) paradigms, which typically rely on two continuously competing speakers. The authors argue that this setup is unrealistic, as natural conversations involve turn-taking speakers. Tracking rapid attention shifts between individual speakers within a conversation is computationally demanding and often unnecessary for practical applications like neuro-steered hearing aids. Instead, the paper proposes a "conversation-tracking" paradigm, where the AAD algorithm identifies the attended conversation as a whole, ignoring internal turn-taking. This approach reduces the frequency of required attention switches, allowing for more relaxed decision windows and better alignment with real-world listening scenarios. To validate this paradigm, the researchers conducted an EEG experiment with 20 normal-hearing participants. The study simulated a complex restaurant environment with three simultaneous two-speaker conversations presented from different spatial locations. The experimental design included four conditions: a reference condition with two competing speakers, a condition with two simultaneous conversations, and two conditions involving three simultaneous conversations (one with sustained attention to a single conversation and one with an intra-trial attention switch). EEG data were recorded using a 64-channel system and processed using a backward least-squares linear decoder to reconstruct speech envelopes. Performance was evaluated using both raw accuracy and an adjusted accuracy metric to account for differing chance levels across conditions. Additionally, participants completed behavioral questionnaires assessing listening effort and speech intelligibility, as well as a Flemish matrix sentence test to measure objective speech-in-noise performance. The results demonstrated that AAD performed significantly above chance in all conditions. There was no significant difference in decoding accuracy between the traditional two-speaker paradigm and the two-conversation paradigm, indicating that conversations can be decoded as effectively as individual speakers. Furthermore, the presence of a third conversation did not significantly impair performance relative to its increased difficulty; adjusted accuracies remained comparable across all conditions. The direction of attention and the presence of intra-trial attention switches also did not significantly affect decoding performance. Behavioral analysis revealed significant correlations between AAD accuracy and self-reported speech intelligibility and listening effort. However, no significant correlation was found between AAD performance and objective speech-in-noise test scores. The study concludes that the conversation-tracking paradigm is a viable and more realistic framework for evaluating AAD algorithms. By focusing on conversations rather than individual speakers, the method maintains high decoding accuracy in complex, multi-source environments while reducing the computational burden associated with tracking rapid speaker switches. These findings support the potential for deploying AAD in practical applications, such as hearing aids, where robust performance in naturalistic, turn-taking conversational settings is essential. The lack of correlation with objective speech intelligibility tests suggests that AAD performance reflects distinct neural mechanisms related to attentional selection rather than general auditory processing ability.
Provenance
The full processing record for this entry. Every stage of this paper's journey through the pipeline is logged — what ran, with which tool and model, how many attempts it took, and when it last completed.
| Stage | Outcome | Tool | Model | Prompt | Attempts | Completed |
|---|---|---|---|---|---|---|
| discover | success | Crossref | — | — | 1 | 2026-08-09 |
| archive | success | canonical_url | — | — | 1 | 2026-08-09 |
| extract | success | pdftotext | — | — | 4 | 2026-08-10 |
| clean | success | clean | — | — | 2 | 2026-08-10 |
| chunk | success | chunk | — | — | 2 | 2026-08-10 |
| embed | success | embed | Qwen/Qwen3-Embedding-8B | — | 2 | 2026-08-10 |
| promote | success | — | — | — | 1 | 2026-08-09 |
| summarize | success | llm | qwen3.6-27b-nvidia | summ-v5 | 2 | 2026-08-10 |
| tag | success | vector_similarity | — | — | 16 | 2026-08-11 |
| verify | success | — | — | — | 2 | 2026-08-10 |
Summary generated by qwen3.6-27b-nvidia on 2026-08-10; verification: verified.
Topics
Ranked by relevance to this paper. Hover a topic for its definition.