N-back Temporal Stability: The Auditory N-back Task as an Unstable Measurement Standard

Wheatley, Camille L.; Cooper, Joel M.; Crabtree, Kaedyn W.; Wise, Ashleigh V. T.; Motzkus, Conner J.; Castro, Spencer C.; Strayer, David L. · 2019 · Wheatley CL, Cooper JM, Crabtree KW, et al.

archive: archived pipeline: cataloged verified

Get this paper ↗ (search — opens at the source; we link to it, we don't host it)

Summary

This study investigates the temporal stability of the auditory N-back task, a standardized reference tool widely used in driving research to benchmark cognitive workload. The authors argue that if a measurement standard drifts over time, as the International Prototype Kilogram did, all data referenced against it become biased. While the N-back is specified in ISO 14198 as a cognitive calibration task, its reliability across repeated exposures in multi-session studies had not been thoroughly verified. The research aims to determine whether N-back performance and associated cognitive demand remain stable or change with practice, and to identify the mechanisms behind any observed changes. Experiment 1 analyzed data from 10 participants who completed an average of 28 on-road driving sessions, with six specific sessions selected for analysis to track changes over 26 exposures. Participants performed the auditory 2-back task while driving, with cognitive workload measured via the Detection Response Task (DRT) and subjective ratings using the NASA-TLX. Results showed that N-back accuracy significantly improved with repeated exposure, approaching a ceiling effect. Concurrently, DRT reaction times decreased, DRT hit rates increased, and subjective workload ratings declined. These changes were specific to the N-back; a single-task baseline remained stable, indicating that the reduction in cognitive demand was not due to general familiarity with the study protocol. Experiment 2 sought to determine whether the improvement in Experiment 1 resulted from participants memorizing the specific digit sequences in the audio files. Twenty participants, who had previously completed at least 10 N-back sessions, performed the task using both the original "Old" sequences and newly generated "New" sequences. Performance metrics, including accuracy, DRT reaction time, hit rate, and subjective workload, showed no significant differences between the Old and New conditions. Post-experiment questionnaires confirmed that participants had not memorized the sequences. This equivalence suggests that the improvement stems from general skill acquisition or strategy adoption, such as subvocal rehearsal, rather than stimulus-specific learning. The findings demonstrate that the auditory N-back is an unstable measurement standard over repeated use. As participants gain experience, the task imposes less cognitive demand, potentially leading to artificially deflated workload estimates in later sessions of multi-session studies. This instability is particularly problematic for automated driving research, where the N-back is frequently used to manipulate cognitive load across multiple trials. The authors recommend limiting repeated N-back exposure, including session number as a covariate in statistical models, reporting participants' exposure history, and considering alternative stable reference tasks like the Surrogate Reference Task (SuRT) for longitudinal designs.

Key finding

The auditory N-back task is NOT stable over repeated use: accuracy increases toward ceiling and cognitive workload decreases significantly across 26 sessions. Improvement transfers to novel digit sequences (general strategy acquisition, not sequence-specific learning). Likely driven by subvocal rehearsal strategy adoption and/or automatization of component processes. Has implications for multi-session studies using N-back as a cognitive reference task — workload estimates will be systematically biased downward in later sessions.

Methodology

on_road

Sample size: Exp 1: N=10 (5F), Mage=25.4; Exp 2: N=20 (10F), Mage=26.5. Both from larger IVIS evaluation (Strayer et al., 2017).

Provenance

The full processing record for this entry. Every stage of this paper's journey through the pipeline is logged — what ran, with which tool and model, how many attempts it took, and when it last completed. Discovered via tag_papers on 2026-05-30 (4 acquisition events logged).

StageOutcomeToolModelPromptAttemptsCompleted
discover success 1 2026-05-07
archive failed pmc 8 2026-06-04
extract success cached 5 2026-08-10
clean success clean 2 2026-08-10
chunk success chunk 2 2026-08-10
embed success embed Qwen/Qwen3-Embedding-8B 2 2026-08-10
enrich success 1 2026-05-07
promote success 2 2026-06-06
summarize success llm qwen3.6-27b-nvidia summ-v5 4 2026-08-10
tag success vector_similarity 27 2026-08-11
verify success 4 2026-08-11

Summary generated by qwen3.6-27b-nvidia on 2026-08-10; verification: verified.

Topics

Ranked by relevance to this paper. Hover a topic for its definition.

Information type

What kind of knowledge this paper contributes, grouped by family — independent of topic (what it is about) and method (how it was studied).