Human Autonomy Teaming and AI Metacognition in Maritime Threat Assessment

Schulze, Kathryn; Gallant, Adele; S Paul, Tanya; Chamberland, Cindy; Lafond, Daniel; Tremblay, Sebastien; Neyedli, Heather · 2026 · Crossref

DOI: 10.54941/ahfe1007173

archive: archived pipeline: cataloged verified

Get this paper ↗ (DOI — opens at the source; we link to it, we don't host it)

Summary

This study establishes baseline human performance data for a simulated maritime threat assessment task, designed to support future research on Human Autonomy Teaming (HAT) and AI metacognition. The research addresses the need for artificial agents to engage in adaptive teamwork processes, such as transparency and self-monitoring, rather than merely performing taskwork. Specifically, the authors aim to validate a simulation environment that will eventually integrate "Cognitive Shadow" (CS), an AI decision-support system capable of modeling expert strategies and estimating its own reliability. The primary objective was to determine if three distinct scenarios, derived from a common template, elicited comparable levels of workload, situation awareness (SA), and self-confidence, thereby ensuring that future improvements attributed to AI integration are not confounded by differences in task difficulty. Thirty-five participants completed one training scenario and three counterbalanced experimental scenarios, each lasting 12 minutes. The task required continuous monitoring and classification of maritime entities as friendly, uncertain, or suspect based on dynamic features such as Automatic Identification System status and intelligence reports. Participants applied predefined decision rules from memory during experimental trials. Data collection included psychometric measures of perceived workload (NASA-TLX), self-confidence, and SA (QUASA), as well as performance metrics including classification accuracy, precision, recall, and confusion matrices. Statistical analyses employed repeated-measures ANOVAs to compare measures across scenarios and scenario positions. Results indicated no significant differences in workload, self-confidence, or SA across the scenarios, confirming the experimental design’s consistency in cognitive demand. However, classification accuracy varied significantly, with Scenario C yielding the highest accuracy (84.92%) and Scenario B the lowest (76.97%). Confusion matrices revealed stable error patterns: the "uncertain" category was most frequently misclassified, and participants exhibited a conservative decision strategy, often downgrading suspect entities to uncertain or uncertain to friendly. This behavior aligns with the Recognition-Primed Decision framework, where operators default to low-risk interpretations to minimize high-cost errors. No systematic learning or fatigue effects were observed across scenario positions, though workload was significantly lower during the training phase due to visual aids. The findings confirm that the simulation provides a controlled, sensitive platform for evaluating HAT interventions. The consistent conservative bias and room for performance improvement offer a clear baseline for assessing the impact of CS. Future phases will introduce AI recommendations and metacognitive confidence displays to evaluate whether these features reduce workload, maintain SA, and foster calibrated trust. This work contributes to the field by demonstrating how baseline human heuristics can be leveraged to design AI teammates that support, rather than replace, human judgment in dynamic, uncertain environments.

Provenance

The full processing record for this entry. Every stage of this paper's journey through the pipeline is logged — what ran, with which tool and model, how many attempts it took, and when it last completed.

StageOutcomeToolModelPromptAttemptsCompleted
discover success Crossref 1 2026-08-09
archive success canonical_url 1 2026-08-09
extract success pdftotext 4 2026-08-10
clean success clean 2 2026-08-10
chunk success chunk 2 2026-08-10
embed success embed Qwen/Qwen3-Embedding-8B 2 2026-08-10
promote success 1 2026-08-09
summarize success llm qwen3.6-27b-nvidia summ-v5 2 2026-08-10
tag success vector_similarity 17 2026-08-11
verify success 2 2026-08-10

Summary generated by qwen3.6-27b-nvidia on 2026-08-10; verification: verified.

Topics

Ranked by relevance to this paper. Hover a topic for its definition.

Information type

What kind of knowledge this paper contributes, grouped by family — independent of topic (what it is about) and method (how it was studied).