Multimodal Large Language Models and Dental Students in Radiographic Landmark Identification: A Novel Grid-Based Assessment
Mehmet Egemen Aydemir, Burcu Sayin, Burcu Yeliz Kollayan, Can Günel, Kadir Cem, Mehmet Ali GülAbstract
Objectives
To introduce a novel grid-based coordinate framework for evaluating spatial landmark localisation competence in multimodal large language models (MLLMs) applied to dental radiography, and to benchmark two MLLMs against dental student consensus.
Methods
GPT-5.4 and Gemini 3.1 Pro were evaluated on 900 queries spanning 200 dental radiographs (100 panoramic, 50 periapical, 50 cephalometric) and 12 landmarks (9 point, 3 area) under zero-shot and guided prompting across three repetitions (10,800 API calls). Performance was scored against a two-rater adjudicated ground truth and a team-adjudicated fourth-year dental student consensus (n = 40). Metrics included Euclidean distance, the successful detection rate, and the Jaccard index.
Results
GPT-5.4 matched student-level performance on cephalometric landmarks but was substantially outperformed on periapical and panoramic tasks. Gemini 3.1 Pro outperformed GPT-5.4 on panoramic and periapical landmarks, while GPT-5.4 retained an advantage on cephalometric points. Guided prompting produced heterogeneous effects, improving localisation of some landmarks while causing severe regression on others; most notably, GPT-5.4 directed the majority of lower-left canine apex predictions to the incorrect lower-right molar region, a prompt-resistant mislocalisation pattern absent in Gemini 3.1 Pro. Both models remained inferior to the student consensus on all periapical and panoramic comparisons, with Gemini 3.1 Pro approaching parity only on cephalometric points under guided prompting.
Conclusion
Current MLLMs demonstrate modality-specific spatial competence insufficient for autonomous clinical deployment. The grid-based framework provides a reproducible, modality-agnostic benchmark for tracking MLLM spatial reasoning across future model generations.