Effects of System Reliability on Workload and Performance in Image Recognition Tasks
DOI: 10.54941/ahfe1004417
archive: archived pipeline: cataloged verified
Get this paper ↗ (DOI — opens at the source; we link to it, we don't host it)
Summary
This study investigates the impact of autonomous system reliability on human operator performance and mental workload within human-machine collaborative systems. Motivated by the widespread integration of autonomy in fields such as unmanned aerial vehicle operations, the research addresses the critical need to understand how imperfect autonomous aids affect operators who must verify system outputs. Specifically, the authors examine how varying levels of autonomy reliability influence task completion time, recognition accuracy, and subjective workload during a vehicle type recognition task. The study aims to determine the reliability threshold at which autonomy becomes beneficial rather than disruptive, providing insights for the design of assistive systems and the allocation of human-machine functions. The experiment employed a within-subjects design with 36 university students tasked with classifying images of vehicles as either passenger or non-passenger. The independent variable was the reliability of the autonomous assistance, set at four levels: 0% (no autonomy), 50%, 70%, and 90%. In conditions with autonomy, the system pre-filtered images, labeling them as passenger vehicles; however, due to imperfect reliability, non-passenger vehicles were occasionally presented. Participants could either accept the system’s judgment or re-evaluate the image independently. Task performance was measured by completion time and accuracy, while mental workload was assessed using the NASA-TLX scale. Data were analyzed using repeated measures ANOVA and linear regression to determine the relationship between reliability levels and operator metrics. The results indicated no significant difference in recognition accuracy across all conditions, with all groups maintaining accuracy above 95%. However, task completion time and subjective workload varied significantly. Autonomy with 90% reliability significantly reduced task completion time and lowered NASA-TLX scores, particularly in mental demand and effort, compared to the no-autonomy baseline. In contrast, 50% reliability tended to increase workload and completion time, though these differences were not statistically significant. Linear regression analysis revealed a negative correlation between reliability and both completion time and workload, identifying a threshold of approximately 55% reliability where autonomy transitions from having a disruptive to a neutral or assistive effect. This threshold is notably lower than the 70% benchmark suggested in previous literature. The authors conclude that autonomy reliability influences operators by altering their task completion strategies, specifically shifting them toward an "all-or-none" approach where high reliability leads to greater reliance on the system, thereby speeding up processing without improving accuracy. Low-reliability autonomy introduces unnecessary cognitive load, as operators cannot fully ignore the system’s presence. These findings underscore the importance of ensuring high reliability in autonomous aids to prevent performance degradation and excessive mental workload. The study suggests that future research should incorporate physiological metrics and longer task durations to better understand strategy stabilization and refine reliability thresholds for real-world applications.
Provenance
The full processing record for this entry. Every stage of this paper's journey through the pipeline is logged — what ran, with which tool and model, how many attempts it took, and when it last completed.
| Stage | Outcome | Tool | Model | Prompt | Attempts | Completed |
|---|---|---|---|---|---|---|
| discover | success | Crossref | — | — | 1 | 2026-08-09 |
| archive | success | canonical_url | — | — | 1 | 2026-08-09 |
| extract | success | pdftotext | — | — | 4 | 2026-08-10 |
| clean | success | clean | — | — | 2 | 2026-08-10 |
| chunk | success | chunk | — | — | 2 | 2026-08-10 |
| embed | success | embed | Qwen/Qwen3-Embedding-8B | — | 2 | 2026-08-10 |
| promote | success | — | — | — | 1 | 2026-08-09 |
| summarize | success | llm | qwen3.6-27b-nvidia | summ-v5 | 2 | 2026-08-10 |
| tag | success | vector_similarity | — | — | 17 | 2026-08-11 |
| verify | success | — | — | — | 2 | 2026-08-10 |
Summary generated by qwen3.6-27b-nvidia on 2026-08-10; verification: verified.
Topics
Ranked by relevance to this paper. Hover a topic for its definition.
Information type
What kind of knowledge this paper contributes, grouped by family — independent of topic (what it is about) and method (how it was studied).
- Empirical Findings: self report data