DOI: 10.1177/10711813261487961 ISSN: 1071-1813

Comparing Human and AI Coding of Situation Awareness: A Signal Detection Analysis of Free-Response Data

Jacqueline Hannan, Federico Puerta Martinez, Noor Dirini, Robina Matyal, Dario Winterton

Large language models (LLMs) are increasingly used for analysis tasks, yet their performance in qualitative review relative to human coding remains unclear. This study compared human and AI coding of free-response data collected from 24 clinical professionals following a medical simulation. Responses were evaluated against 23 expert-derived situational elements by human coders and four LLMs: ChatGPT 5.2, Claude 4.5 Sonnet, Microsoft Copilot for Work, and Qwen3-32b. Agreement was assessed using Cohen’s kappa, signal detection analysis, and Spearman correlations of participant-level scores. All models demonstrated limited item-level agreement with human coding and showed systematic conservative bias characterized by significantly more false negatives than false positives ( p  < .001). Despite weak event-level agreement, participant-level correlations with human-derived situation awareness scores remained moderate ( r  = .648–.793). Cross-model interpretation was limited by differences in deployment configuration and parameter control. Findings suggest that current LLMs may support aggregate-level qualitative assessment but remain insufficient substitutes for detailed human qualitative analysis.