DOI: 10.24938/kutfd.1922944 ISSN: 1302-3314

SCENARIO-BASED PREHOSPITAL TRIAGE OF CARBON MONOXIDE POISONING: COMPARING FIVE LARGE LANGUAGE MODELS WITH EMERGENCY MEDICAL SERVICES PERSONNEL IN A PROSPECTIVE STUDY

Vildan Özer, Özlem Bülbül, Efnan Bayrak Erbolukbas, Serdar Karakullukçu, Aynur Şahin
Objective: Carbon monoxide (CO) poisoning requires rapid identification and timely decisions regarding the need for hyperbaric oxygen therapy (HBOT) to improve clinical outcomes. This study aimed to compare the decision-making performance of emergency medical services (EMS) personnel and large language models (LLMs) in accurately determining the need for HBOT during the prehospital phase of CO poisoning cases, and to explore the potential implementation of LLMs as decision-support tools for patient triage and referral.Material and Methods: In this prospective, simulation-based diagnostic accuracy study, 128 standardized scenario-based clinical cases (64 requiring HBOT, 64 not requiring HBOT) were developed based on established indications. Sixty EMS personnel and five LLMs (GPT-4o, GPT-4.5, GPT-o3, Gemini 2.5 Pro, and DeepSeek-R1) evaluated each scenario. Their responses were compared with the gold standard answers established by consensus between a medical toxicologist and a hyperbaric medicine specialist.Results: The overall accuracy of the EMS personnel was 65.1%, whereas the highest accuracy was observed with GPT-o3 (96.9%). All LLMs demonstrated 100% sensitivity, although the specificity varied, ranging from 26.6% (GPT-4.5) to 93.8% (GPT-o3). The accuracy of GPT-o3 (96.9%, p

More from our Archive