Performance and workload trends: The effects of repeated exposure to "high" demand tasks

Wheatley, CL; Esplin, J; Loveless, SM; Cooper, JM; Biondi, F; Strayer, DL · 2018 · publications_jsonl

DOI: 10.1177/1541931218621002

archive: archived pipeline: cataloged verified

Get this paper ↗ (DOI — opens at the source; we link to it, we don't host it)

Summary

This study investigates the stability of performance and workload metrics for two standardized reference tasks—the auditory N-back and the Surrogate Reference Task (SuRT)—when participants are repeatedly exposed to them. These tasks are widely used as high-demand benchmarks to calibrate the cognitive and visual workload of in-vehicle information systems. The research addresses a critical gap: if repeated exposure improves performance or reduces perceived workload, these tasks may lose their validity as consistent calibration tools. The authors hypothesized that changes in task performance over time could negate their effectiveness as stable reference points for comparing secondary driving tasks. The study utilized data from a larger collaboration between the University of Utah and the AAA Foundation for Traffic Safety, focusing on ten participants who completed at least 26 driving sessions. Data were collected on-road in physical vehicles on a residential route. Participants performed the 2-back variant of the N-back task, which assesses working memory, and the SuRT, a visual-manual task requiring identification of a target circle among distractors. Workload was measured objectively using a Detection Response Task (DRT), where reaction time to a vibrotactor indicated cognitive load and hit rate to a head-up display light indicated visual load. Subjective workload was assessed using the NASA Task Load Index (TLX). One-way repeated measures ANOVAs compared performance and workload metrics across six levels of exposure. Results indicated divergent trends for the two tasks. For the N-back, performance accuracy improved significantly with repeated exposure, approaching a ceiling effect. Concurrently, DRT reaction times decreased, and subjective TLX scores dropped, indicating a significant reduction in cognitive workload. This suggests the N-back becomes less cognitively demanding over time, limiting its utility as a stable high-demand benchmark for repeated exposures. In contrast, SuRT performance also improved, with participants completing more trials per minute. However, DRT hit rates remained stable, indicating that visual demand did not decrease despite performance gains. Although subjective TLX scores for the SuRT decreased, the objective visual workload remained constant, validating the SuRT’s stability as a high visual demand benchmark. The findings imply that the N-back is an unstable reference task for cognitive workload in studies involving repeated exposures, as participants likely develop strategies or improve working memory capacity, reducing the task's demand. Researchers should exercise caution when using the N-back in longitudinal designs or consider modifications, such as varying sequences, to maintain its difficulty. Conversely, the SuRT remains a valid and stable benchmark for visual demand, as its structural interference with driving attention persists regardless of performance improvements. The study highlights the importance of validating reference tasks over time to ensure accurate assessment of driver workload in human factors research.

Key finding

Repeated exposure erodes the N-back's value as a high-cognitive-demand benchmark (workload drops as performance improves) but leaves the SuRT's visual-demand calibration intact.

Methodology

other

Sample size: 10

Provenance

The full processing record for this entry. Every stage of this paper's journey through the pipeline is logged — what ran, with which tool and model, how many attempts it took, and when it last completed. Discovered via author_sweep_intake on 2026-05-28 (4 acquisition events logged).

StageOutcomeToolModelPromptAttemptsCompleted
discover success author_sweep 2 2026-05-28
archive failed pmc 8 2026-06-04
extract success cached 5 2026-08-10
clean success clean 2 2026-08-10
chunk success chunk 2 2026-08-10
embed success embed Qwen/Qwen3-Embedding-8B 2 2026-08-10
enrich success semantic_scholar 1 2026-06-04
promote success 2 2026-06-06
summarize success llm qwen3.6-27b-nvidia summ-v5 4 2026-08-10
tag success vector_similarity 27 2026-08-11
verify success 4 2026-08-11

Summary generated by qwen3.6-27b-nvidia on 2026-08-10; verification: verified.

Topics

Ranked by relevance to this paper. Hover a topic for its definition.

Information type

What kind of knowledge this paper contributes, grouped by family — independent of topic (what it is about) and method (how it was studied).