DOI: 10.7126/cumudj.1794852 ISSN: 1302-5805

Comparative Evaluation of Large Language Models’ Accuracy in Answering Multiple-Choice Restorative Dentistry Questions From a National Specialty Examination

Çilem Bulut, Gülben Çolak, Gürkan Çolak
Objective: Although the integration of large language models (LLMs) into dental education is rapidly increasing, their actual performance in domain-specific assessments remains unclear. This study aimed to evaluate and compare the accuracy of four LLMs (ChatGPT-4.0, Gemini Advanced 1.5 Pro, DeepSeek-V3, and Perplexity) on restorative dentistry questions in a dental specialty examination. Materials and methods: A total of 127 multiple-choice questions from the Turkish Dental Specialty Examination (DUS) conducted between 2012 and 2021 were collected and categorized into 19 content areas. Each question was entered into LLMs in Turkish with standardized instructions. Responses were recorded, and their accuracy was determined according to official answer keys. Statistical differences were analyzed using the appropriate tests. Results: ChatGPT-4.0 had the highest accuracy rate (93.65%), followed by Gemini (82.54%), DeepSeek (71.43%), and Perplexity (65.87%). Significant differences were observed between ChatGPT and DeepSeek (p = 0.027) and Perplexity (p = 0.004), but not between ChatGPT and Gemini (p = 0.118). All models showed higher accuracy in theoretical questions but lower performance in clinically oriented areas such as bleaching and cavity preparation. Conclusion: ChatGPT-4.0 demonstrated the highest overall accuracy among the language models evaluated and shows promise as a supportive tool in theoretical dental education. However, its limited performance in clinical domains underlines the need for careful implementation under physician supervision.