Human–AI Collaboration in Automated X-Ray Screening: Effects of Alarm Types and Reliability Levels on Operator Performance in Subway Security
DOI: 10.54941/ahfe1007456
archive: archived pipeline: cataloged
Get this paper ↗ (DOI — opens at the source; we link to it, we don't host it)
Summary
This study investigates how the format of automated diagnostic advice should be matched to the reliability of an automated diagnostic aid system (ADAS) in subway X-ray baggage screening. The research addresses a critical gap in human–AI collaboration: while ADASs reduce operator workload, the optimal alarm design depends on system accuracy, and it remains unclear how operators’ performance and trust vary when alarm transparency is adjusted across different reliability levels. The experiment employed a 3 × 3 mixed factorial design with 18 graduate students (six per reliability group). Alarm type (binary, likelihood, automated decision) was manipulated within subjects, while ADAS reliability (70%, 80%, 90%) was manipulated between subjects. Participants completed 100 trials per alarm type in a simulated screening task using colored X-ray images with a 30% target prevalence. Performance was measured via signal detection sensitivity (d′) and response times, while subjective trust was assessed across five dimensions using a questionnaire. Statistical analysis included two-way mixed ANOVAs for objective measures and Kruskal–Wallis tests for trust ratings. Results revealed a significant alarm type × reliability interaction on sensitivity (F(2, 15) = 7.05, p = .007, ηp² = .48), with no significant main effects. Under the binary alarm, operator sensitivity increased monotonically with reliability. In contrast, under the likelihood alarm, sensitivity peaked at 80% reliability but deteriorated at 90%, falling below the ADAS’s standalone performance. This suggests that graded alarms provide beneficial diagnostic information when system reliability is low, but become counterproductive when the system is highly accurate, as the additional ambiguity increases cognitive load without adding value. Subjective trust patterns aligned with these objective findings: participants preferred the likelihood alarm at lower reliability levels but favored the automated decision mode at 90% reliability, where they expressed the highest trust. The added value of the human operator diminished as system reliability increased, with joint performance gains over the ADAS alone shrinking from approximately 70% at 70% reliability to 11% at 90% reliability. The findings indicate that alarm transparency should be adapted to automation reliability to optimize human–machine teaming. For low-reliability systems, likelihood alarms are preferable as they recruit operator attention and support informed overrides. For high-reliability systems, binary alarms or automated decision modes are superior, minimizing unnecessary decisional load and enabling efficient triage. These results provide actionable design guidelines for safety-critical screening operations, emphasizing that trust calibration is contingent on the form in which automation communicates its assessments. Limitations include the small sample size, use of student participants rather than professional screeners, and the high target prevalence used in the simulation, which may limit generalizability to field settings.
Provenance
The full processing record for this entry. Every stage of this paper's journey through the pipeline is logged — what ran, with which tool and model, how many attempts it took, and when it last completed.
| Stage | Outcome | Tool | Model | Prompt | Attempts | Completed |
|---|---|---|---|---|---|---|
| discover | success | Crossref | — | — | 1 | 2026-08-09 |
| archive | success | canonical_url | — | — | 1 | 2026-08-09 |
| extract | success | cached | — | — | 5 | 2026-08-23 |
| clean | success | clean | — | — | 2 | 2026-08-10 |
| chunk | success | chunk | — | — | 2 | 2026-08-10 |
| embed | success | embed | Qwen/Qwen3-Embedding-8B | — | 2 | 2026-08-10 |
| promote | success | — | — | — | 1 | 2026-08-09 |
| summarize | success | llm | qwen3.8-27b-gittensor | summ-v5 | 3 | 2026-08-23 |
| tag | success | vector_similarity | — | — | 17 | 2026-08-11 |
| verify | success | — | — | — | 1 | 2026-08-09 |
Summary generated by qwen3.8-27b-gittensor on 2026-08-23; verification: pending re-verification.
Topics
Ranked by relevance to this paper. Hover a topic for its definition.