Comparing Human and AI Coding of Situation Awareness: A Signal Detection Analysis of Free-Response Data
Jacqueline Hannan, Federico Puerta Martinez, Noor Dirini, Robina Matyal, Dario Winterton
Large language models (LLMs) are increasingly used for analysis tasks, yet their performance in qualitative review relative to human coding remains unclear. This study compared human and AI coding of free-response data collected from 24 clinical professionals following a medical simulation. Responses were evaluated against 23 expert-derived situational elements by human coders and four LLMs: ChatGPT 5.2, Claude 4.5 Sonnet, Microsoft Copilot for Work, and Qwen3-32b. Agreement was assessed using Cohen’s kappa, signal detection analysis, and Spearman correlations of participant-level scores. All models demonstrated limited item-level agreement with human coding and showed systematic conservative bias characterized by significantly more false negatives than false positives (