The power and sensitivity of four core driver workload measures for benchmarking the distraction potential of new driver vehicle interfaces

McDonnell, AS; Imberger, K; Poulter, C; Cooper, JM · 2021 · publications_jsonl

DOI: 10.1016/j.trf.2021.09.019

archive: archived pipeline: cataloged verified

Get this paper ↗ (DOI — opens at the source; we link to it, we don't host it)

Summary

This study evaluates the statistical power and sensitivity of four core driver workload measures to determine their utility in benchmarking the distraction potential of in-vehicle human-machine interfaces (HMIs). As in-vehicle technology increases, assessing how secondary tasks compete for visual, manual, and cognitive resources is critical for safety standards and New Car Assessment Program (NCAP) testing. The research addresses the need for valid, predictive workload metrics that can detect differences in distraction potential across various vehicle systems. Specifically, the authors investigate which measures are most sensitive to workload variations and how these measures correlate with one another, aiming to refine evaluation protocols for future vehicle safety assessments. The analysis utilized a secondary dataset from 173 participants who drove 40 different vehicles equipped with 50 distinct IVIS configurations. Participants completed secondary tasks involving text messaging, navigation, audio entertainment, and calling via auditory vocal, center stack, or center console modalities. Four workload metrics were collected: DRT Reaction Time (cognitive demand), DRT Miss Rate (visual demand), NASA-TLX (subjective workload), and Task Interaction Time (temporal demand). The study employed variance and power analyses to determine the sample sizes required to detect significant differences in workload between vehicles and to assess the interrelationships among the four measures. The results demonstrated that Task Interaction Time was the most sensitive measure for detecting differences in driver workload between different HMIs, followed by DRT Miss Rate, NASA-TLX, and finally DRT Reaction Time. Power analyses indicated that measures with lower sensitivity, such as DRT Reaction Time, would require significantly larger sample sizes to achieve statistical significance. Furthermore, correlations between the four measures were relatively weak, suggesting that each metric captures a unique aspect of driver workload rather than redundant information. These findings imply that Task Interaction Time, combined with a reliable visual demand metric like DRT Miss Rate, offers a more efficient and sensitive approach to evaluating HMI distraction potential than relying solely on DRT Reaction Time or subjective NASA-TLX scores. The study concludes that a multi-metric approach is necessary for comprehensive distraction assessment, but prioritizing temporal and visual demand measures can improve the validity and efficiency of future vehicle safety testing regimes. This work provides empirical guidance for selecting workload metrics that best discriminate between varying levels of interface complexity and driver burden.

Key finding

Task Interaction Time was the most sensitive measure of between-vehicle workload differences, followed by DRT Miss Rate, NASA-TLX, and finally DRT Reaction Time. DRT Miss Rate variance attributable to vehicle ranged from 10% (collapsed) to 20% (navigation), with 80%-power sample size requirements of 10-17 per task. DRT Reaction Time accounted for only ~5% of variance overall (n~89 needed), suggesting cognitive load operated in an on-off fashion at a similar elevated level across vehicles, tasks, and modalities. Correlations between the four measures were weak, indicating each captures a partially unique aspect of workload; Task Interaction Time was largely independent of the other three.

Methodology

on_road

Sample size: N=173 drivers across 40 vehicles and 50 IVIS configurations

Provenance

The full processing record for this entry. Every stage of this paper's journey through the pipeline is logged — what ran, with which tool and model, how many attempts it took, and when it last completed. Discovered via gog_drive on 2026-06-06 (2 acquisition events logged).

StageOutcomeToolModelPromptAttemptsCompleted
discover success 1 2026-05-06
archive failed pmc 12 2026-06-04
extract success cached 5 2026-08-10
clean success clean 2 2026-08-10
chunk success chunk 2 2026-08-10
embed success embed Qwen/Qwen3-Embedding-8B 2 2026-08-10
enrich failed 5 2026-08-07
promote success 2 2026-06-06
summarize success llm qwen3.6-27b-nvidia summ-v5 4 2026-08-10
tag success vector_similarity 26 2026-08-11
verify success 4 2026-08-11

Summary generated by qwen3.6-27b-nvidia on 2026-08-10; verification: verified.

Topics

Ranked by relevance to this paper. Hover a topic for its definition.

Information type

What kind of knowledge this paper contributes, grouped by family — independent of topic (what it is about) and method (how it was studied).