DOI: 10.1097/md.0000000000050122 ISSN: 0025-7974
Artificial intelligence performance in the emergency medicine subspecialty examination conducted in Türkiye
Merve Ağaçkiran, Sinan Önder, Ümit Can Yürekli, İlter Ağaçkiran
Emergency medicine specialists often pursue subspecialty training worldwide. In Türkiye, subspecialization in critical care medicine was introduced in March 2024, with the first entrance examination for subspecialty training in medicine (YDUS) examination having been conducted on December 15, 2024 by the Measurement, Selection, and Placement Center. Medical applications of artificial intelligence (AI), particularly GPT-4 Omni (GPT-4o), GPT-4, and Gemini-Advanced, have garnered considerable attention. This study aimed to evaluate the performance of these AI models in answering emergency medicine YDUS questions, marking the first assessment of the role of AI in this examination. The performance of 3 AI models (GPT-4, GPT-4o, and Gemini-Advanced) on questions from the emergency medicine YDUS examination was evaluated. The examination included 60 multiple-choice questions, of which 10% were publicly available. Questions were classified as clinical or factual. Responses of the AI models were analyzed using Cochran
Q
test as the omnibus test, with exact Bonferroni-adjusted McNemar tests for pairwise comparisons where applicable. No significant differences in the correct responses for both clinical and factual questions were observed between each AI model (
P
values: GPT-4o, 1.000; Gemini-Advanced, .554; and GPT-4, 1.000). GPT-4o significantly outperformed Gemini-Advanced in clinical (92.6% vs 70.4%) and factual questions (90.9% vs 78.8%) (
P
values: clinical, .021; factual, .039). A comparison of the overall performance showed an omnibus significant difference (
P
= .001); however, post hoc pairwise comparisons revealed that only GPT-4o (91.7%) significantly outperformed Gemini-Advanced (75%), whereas GPT-4 (88.3%) did not show a statistically significant difference from Gemini-Advanced after adjustment. This study found that both GPT-4o and GPT-4 significantly outperformed Gemini-Advanced in answering Turkish emergency medicine YDUS questions. While GPT-4o achieved the highest numerical accuracy, there was no statistically significant difference between GPT-4o and GPT-4. Both models demonstrated high accuracy in this examination dataset. Although these findings highlight their potential as supplementary learning tools, strong examination performance does not establish clinical readiness or definitive educational usefulness. Gemini-Advanced exhibited weaker performance but frequently advised expert consultation. However, this study did not formally evaluate ethical behavior, safety, or the appropriateness of these refusals.