DOI: 10.31067/acusaglik.1956329 ISSN: 1309-470X

Accuracy and Completeness of Contemporary Large Language Models in Prosthodontics: An Expert-Based Comparative Study

Elif Yiğit İren, Hatice Betül Üçkuyu
Purpose: To compare the scientific accuracy and response completeness of four contemporary LLMs: ChatGPT-5.2, Gemini 3, Copilot, and DeepSeek-V3.2 across major prosthodontic domains and question formats.Methods: Fifty prosthodontic questions covering removable prosthodontics, fixed prosthodontics, implantology, temporomandibular disorders and occlusion, dental materials science were developed from standard textbooks and clinical guidelines. Questions were presented in yes/no, multiple-choice, and open-ended formats to reflect different levels of clinical reasoning. Responses generated by ChatGPT-5.2 (OpenAI), Gemini 3 (Google), Microsoft Copilot (Microsoft), and DeepSeek-V3.2 were independently evaluated by five prosthodontists using Likert-based scales for accuracy and completeness. Inter-rater reliability was assessed to ensure scoring consistency, and model performances were compared across domains and question types using nonparametric statistical tests.Results: Overall accuracy differed significantly among the four models (p0.05), although descriptive trends suggested slightly higher scores for yes/no questions and lower scores for multiple-choice items. Response completeness also varied significantly across models (p0.05).Conclusion: Contemporary large language models demonstrate acceptable theoretical performance in prosthodontics; however, clinically relevant differences persist in response completeness and consistency. These models may support prosthodontic education and preliminary information retrieval, but cannot replace expert clinical judgment.

More from our Archive