Temporal Reproducibility of Fracture Interpretation in Forensic Radiography: A Multispecialty Comparison of Physicians and Vision Language Models Including Fracture Subtype Description
Halit Canberk Aydogan, Adem Köksal, Ali Aygün, Hacer Yaşar Teke, Feyza Nur Çatalbaş, İbrahim ÇaltekinBackground: Accurate fracture interpretation on plain radiographs is critical for both trauma care and medico-legal decision-making, where reproducibility is as important as point accuracy. Although vision language models (VLMs) have shown promising diagnostic performance, their temporal stability in forensic radiography remains unclear. Methods: We analyzed 300 forensic radiographs (150 fracture-positive, 150 fracture-negative) from six long bones, independently evaluated by three emergency medicine physicians, three forensic medicine physicians, and three VLMs (ChatGPT-5.2, Gemini 3 Pro, Claude Sonnet 4.5) using an identical task format. Assessments included fracture presence and structured fracture subtype description (bone, morphology, displacement). VLM evaluations were repeated after one month under identical conditions. Results: Physician accuracy ranged from 79.7% to 98.7%. Emergency physicians reached the higher median sensitivity (92.0% against 76.7%), while specificity among the forensic readers was the more tightly clustered (median 92.0%, range 87.3–100.0%). ChatGPT-5.2 was the most accurate model (83.0%; sensitivity 70.0%, specificity 96.0%), followed by Gemini 3 Pro (77.7%), whereas Claude Sonnet 4.5 reached only 46.3% because of an extreme false-positive tendency (specificity 11.3%). Over one month, accuracy changed by −5.2, −4.1 and +8.6 percentage points, but these net figures concealed considerable case-level movement: within-model agreement ranged from near chance to substantial (mean Cohen κ 0.131 to 0.622), and the F1-score of Claude Sonnet 4.5 fell by 13.1 points despite its higher accuracy. Subtype descriptions were frequently correct once a fracture had been detected, but end-to-end subtype accuracy remained low. Conclusions: Current vision language models demonstrate encouraging diagnostic performance; however, their temporal reproducibility remains inadequate for independent medico-legal fracture interpretation. These findings highlight that reproducibility, in addition to diagnostic accuracy, should be considered a core benchmark when evaluating VLMs for high-stakes clinical and forensic use. Larger multicenter studies using independent external datasets are needed before forensic application is considered.