Assessing ChatGPT‐4o's Potential as a Support Tool in Neuroradiology Cases: A Comparison of Performance by Using Radiologic Images
Levent Akman Solim, Duygu Atasoy, Sabahattin Yuzkan, Vitali Koch, Ahmed Ait Bachir, Ibrahim Yel, Thomas J. VoglABSTRACT
Background
In clinical settings, particularly in radiology, artificial intelligence and large language models have already demonstrated promising results. They have the potential to become an integral component of radiologists' workflows in the future, and neuroradiology is likely to benefit from their increasing integration into clinical practice. This study aims to evaluate Chat Generative Pre‐trained Transformer 4 Omni (ChatGPT‐4o) as a supportive tool in neuroradiology by assessing its ability to generate accurate radiology reports from imaging and patient history and its impact on radiologists' diagnostic accuracy.
Methods
Retrospective analysis was conducted using only radiological images and brief patient admission history from 30 neuroradiology cases published publicly between 2015 and 2024. ChatGPT‐4o generated radiology reports with final and differential diagnoses. One radiologist and one radiology resident independently reviewed cases without and with the reports created by ChatGPT‐4o at a 4‐week interval. Diagnostic accuracy was compared to the published gold‐standard histopathological diagnoses. Also, ChatGPT‐4o was asked to detect the orientation, sequence, and contrast usage. To compare the accuracies, exact McNemar's test and Wilcoxon signed‐rank test with Bonferroni correction was used for statistical analysis.
Results
Evaluation of radiology reports demonstrated an accuracy of final diagnosis from ChatGPT‐4o 33.33% (10/30), experienced radiologist 56.67% (17/30), and radiology resident 23.33% (7/30) ( p < 0.039 for experienced radiologist vs. ChatGPT‐4o and p < 0.375 for resident vs. ChatGPT‐4o). Accuracy rates for differential diagnosis were: ChatGPT‐4o 60.50% ± 21.30%, experienced radiologist 64.83% ± 17.24%, and radiology resident 42.50% ± 16.90% ( p = 0.242 for experienced radiologist vs. ChatGPT‐4o and p = 0.0004 for resident vs. ChatGPT‐4o). Diagnostic performance with the use of ChatGPT‐4o as a support tool demonstrated an accuracy of final diagnosis: experienced radiologist 56.67% (17/30) and radiology resident 30.00% (9/30) ( p = 1.000 and p = 0.500, respectively) and of differential diagnosis: experienced radiologist 66.50% ± 19.30% and radiology resident 49.50% ± 15.56% ( p = 0.500 and p = 0.013, respectively).
Conclusions
ChatGPT‐4o demonstrated limited autonomous and supportive diagnostic accuracy in complex neuroradiology cases. While its assistance did not improve the experienced radiologist's performance and produced only a modest, non‐significant improvement for the resident, a trend toward improved differential diagnosis generation was observed. These findings suggest that multimodal large language models may have limited supportive value, particularly for less experienced clinicians, and require validation in larger studies.