DOI: 10.1111/eje.70268 ISSN: 1396-5883

Assessing the Educational Role of Large Language Models in Dental Training: A Decade‐Long Analysis of Text‐Based and Visual Questions in a National Examination

Sedef Ayse Tasyapan, Didem Özer

ABSTRACT

Background

Large language models (LLMs), including ChatGPT, Gemini and DeepSeek, are increasingly explored as supportive tools in health professions education. However, their educational utility across different knowledge domains and question formats, particularly those involving visual content, remains insufficiently understood.

Objective

This study aimed to evaluate the educational potential of three advanced LLMs in dental training by analysing their performance on a national specialty examination over a 10‐year period, with particular emphasis on domain‐specific accuracy and differences between text‐based and visual questions.

Methods

A total of 1560 multiple‐choice questions from 13 administrations of the Turkish Dental Specialty Examination (DUS) between 2012 and 2021 were included. All questions were translated into English and categorized into nine dental specialties. Each question was individually entered into ChatGPT‐4.0, Gemini Advanced and DeepSeek in isolated sessions. Model responses were compared with official answer keys, and accuracy rates were analysed across years, specialties and question types. Statistical analyses included chi‐square tests, one‐way ANOVA and Kruskal–Wallis tests.

Results

ChatGPT‐4.0 achieved the highest overall accuracy (86.93%), followed by Gemini (82.86%) and DeepSeek (82.34%). Performance varied across specialties, with higher accuracy observed in Basic Sciences and Oral Surgery, and lower performance in Orthodontics and Endodontics. A substantial decrease in accuracy was observed for visual questions (ChatGPT: 52.38%; Gemini: 45.24%; DeepSeek: 2.38%) compared to text‐based items (all models > 83%). ChatGPT demonstrated more stable performance across years, whereas Gemini and DeepSeek showed greater variability.

Conclusions

LLMs demonstrate strong potential as supportive tools in dental education, particularly for text‐based knowledge assessment. However, their limited performance in visual question contexts highlights an important constraint for their integration into image‐dependent domains such as dental diagnostics. These findings underscore the need for cautious and context‐aware implementation of AI tools in dental curricula and assessment practices.

More from our Archive