Information matters, but cues are dangerous: a comparative evaluation of three multimodal AI chatbots in oral and maxillofacial radiology
Chang-ki MinAbstract
Objectives
To compare the diagnostic performance of three multimodal AI chatbots on oral and maxillofacial radiographic images and to examine how additional information, delivery mode, and user-suggested diagnoses affect accuracy.
Methods
Three AI chatbots (GPT-5.1, Gemini 3 Flash, Claude Opus 4.7) were tested on 90 cases comprising normal controls, osteomyelitis, and benign jaw lesions. Inputs were combinations of a panoramic image, a cropped panoramic image, an axial CBCT image or a text-based cue. Inputs were delivered all at once or sequentially. Accuracy was scored at category and specific-diagnosis levels using non-parametric tests with false-discovery-rate correction.
Results
With the panoramic image alone, accuracy for diseased cases was low (0–60%) but rose to as high as 38–92% in each model's best condition with added information. A cropped image was the most consistently beneficial additional visual input, whereas an axial CBCT image provided less improvement. GPT-5.1 recognised normal cases well but missed lesions, Gemini 3 was sensitive but less specific, and Claude 4.7 defaulted to benign diagnoses. Correct verbal cues increased accuracy, whereas a misleading cue caused decline. Gemini 3 accepted a false benign suggestion in 96% of cases it had initially classified as normal. For benign lesions, specific-diagnosis accuracy was almost half of category-level accuracy.
Conclusions
Accuracy was strongly affected by model choice, information type and user-suggested diagnoses, but no model reached a level that would be acceptable for clinical use.
Advances in knowledge
This multi-model comparison isolates the effects of information type, delivery mode, and sycophancy in oral and maxillofacial radiology.