DOI: 10.1055/s-0046-1827461 ISSN: 0971-3026

Patterns of Errors and Hallucinations Among ChatGPT, Perplexity, Qwen, and Copilot in Answering ACR DXIT Radiology Questions

Kurian C. Eapen, Viswajit Krothapalli, Rishabh Vaid, Manoj J. Dhinagar, Lingala J. P. Raju, Rashi Goyal, Aparna Irodi

Abstract

Large language models (LLMs) are increasingly used in radiology workflows. In this study, we aimed to compare the performance and characteristics of errors made by ChatGPT, Perplexity, Qwen, and Microsoft Copilot using standardized radiology questions.

This cross-platform prospective study used 600 American College of Radiology Diagnostic Radiology In-Training Examination questions (360 text-based, 240 image-based). Each question was independently answered by all four LLMs. The questions were categorized by type of question (text vs. image), subspecialty, and difficulty level based on Bloom's taxonomy. Errors were categorized by type of error committed and clinical impact (negligible, mild, major, catastrophic).

Across models, accuracy was higher for text-based (86–91%) than image-based questions (51–64%). All models showed poorer performance with increasing cognitive complexity, with “analysis” questions showing the poorest performance.

Inter-model agreement was fair to moderate for text-based questions (κ range: 0.32–0.56) and moderate to strong for image-based questions (κ range: 0.42–0.67). High agreement was observed for major or catastrophic errors across all models (κ range: 0.813–0.959). The most common reason for error was related to misinterpretation of images, logical reasoning errors, and factual inaccuracies. More than 70% of incorrect responses were delivered with a confident tone.

Similar patterns of hallucinations and high agreement in errors indicate shared structural vulnerabilities across LLM models. Further research into mitigating hallucinations is essential before integration into clinical workflows.

More from our Archive