DOI: 10.7126/cumudj.1953387 ISSN: 1302-5805
Evaluation of large language model responses to different types of questions based on European Society of Endodontology position statements
Merve Defişet, Merve Yeniçeri Özata, Ayşenur Çam Objectives: This study aimed to compare the performance of DeepSeek-V3, ChatGPT 4.0, ChatGPT 4.5, Gemini Advanced 2.0, and Grok in answering single-correct answer multiple-choice (SCQ), multiple-correct answer multiple-choice (MCQ), binary (true/false) (BQ), and open-ended questions (OEQ) based on the European Society of Endodontology (ESE) position statements.Materials and Methods: A total of 70 questions (24 SCQ/MCQ, 29 BQ, and 17 OEQ) were posed to five LLMs. Responses were scored on a 0–10 scale for comprehensiveness, scientific accuracy, clarity, and relevance. SCQ, MCQ, and BQ formats were additionally evaluated for correctness (1 = correct, 0 = incorrect). Data were analyzed using Pearson’s chi-square test and the Kruskal-Wallis H test (p < 0.05).Results: In the SCQ/MCQ format, Gemini Advanced 2.0 scored significantly higher than Grok. In the BQ format, ChatGPT 4.5 and Gemini Advanced 2.0 achieved significantly higher scores than ChatGPT 4.0 and Grok. In the OEQ format, ChatGPT 4.5 obtained significantly higher scores than Grok (p < 0.05). Accuracy rates were 67.9% for ChatGPT 4.5, 66.0% for Gemini Advanced 2.0, 58.5% for DeepSeek-V3, 52.8% for ChatGPT 4.0, and 49.1% for Grok (p > 0.05). In all models, scientific accuracy scores were significantly lower than relevance scores (p < 0.05).Conclusions: ChatGPT 4.5 and Gemini Advanced 2.0 demonstrated stronger performance than the other models, whereas Grok showed lower response-content performance. However, since LLMs may appear convincing even when generating incorrect clinical information, they should be used only as supplementary tools in endodontic education and preliminary information retrieval. Clinically relevant outputs should be verified by an endodontist or an appropriately qualified clinician.
More from our Archive
-
DOI: 10.68381/jca02008 2026
Proximal Smoothness and the Lower-C
2
Property F. H. Clarke, R. J. Stern, P. R. Wolenski
-
DOI: 10.68381/jca13044 2026
Characterizations of Prox-Regular Sets in Uniformly Convex Banach Spaces Frédéric Bernard, Lionel Thibault, Nadia Zlateva
-
DOI: 10.68381/jca15047 2026
Brøndsted-Rockafellar Property and Maximality of Monotone Operators Representable by Convex Functions in Non-Reflexive Banach Spaces Maicon Marques Alves, Benar Fux Svaiter
-
DOI: 10.68381/jca16027 2026
Proximal Smoothness and the Exterior Sphere Condition Chadi Nour, Ron J. Stern, Jean Takche
-
DOI: 10.68381/jca16053 2026
A New Old Class of Maximal Monotone Operators Maicon Marques Alves, Benar Fux Svaiter
-
DOI: 10.68381/jca13045 2026
Maximal Monotonicity via Convex Analysis Jonathan Borwein
-
DOI: 10.68381/jca08009 2026
Variational Inequalities and Regularity Properties of Closed Sets in Hilbert Spaces Giovanni Colombo, Vladimir V. Goncharov
-
DOI: 10.68381/jca17060 2026
Existence and Uniqueness of Solutions for Non-Autonomous Complementarity Dynamical Systems Bernard Brogliato, Lionel Thibault
-
DOI: 10.68381/jca01001 2026
Variational Sum of Monotone Operators H. Attouch, J.-B. Baillon, M. Théra
-
DOI: 10.68381/jca22017 2026
Weak Convexity of Sets and Functions in a Banach Space Grigorii E. Ivanov