DOI: 10.32322/jhsm.1923211 ISSN: 2636-8579

Evaluation of large language model responses to expert questions in anterior implant dentistry: quality, accuracy, and readability

Dalndushe Abdulai, Raghib Suradi, Mehran Moghbel
Aims: This study evaluated the quality, accuracy, and readability of responses generated by four Artificial Intelligence (AI) chatbots based on large language models (LLMs): ChatGPT-5.2, Gemini 3, DeepSeek-V3.2, and Grok-4.1, when responding to expert-generated questions in anterior implant dentistry. The aim was to evaluate their potential role as educational and adjunct informational tools in treatment planning and esthetic zone management, while also examining the readability of the generated responses.Methods: Thirty-six standardized questions covering diagnosis, esthetic risk assessment, implant positioning, surgical planning, peri-implant soft tissue management, esthetic complications, and preventive strategies were developed by three prosthodontists experienced in implant dentistry. Responses from ChatGPT, Gemini, DeepSeek, and Grok were independently evaluated by three experts using the modified DISCERN (mDISCERN) score, Global Quality Score (GQS), and a five-point Accuracy Score. Misinformation and Harm scores were also recorded. Repeated-measures comparisons were performed using the Friedman test with Wilcoxon signed-rank post hoc analysis and Bonferroni correction.Results: ChatGPT demonstrated the highest mean scores for informational reliability (mDISCERN: 3.33±0.72) and accuracy (3.56±0.84); however, differences in accuracy were not statistically significant. Gemini showed the highest numerical GQS value (GQS: 3.75±0.77); however, the overall effect size was small, and no GQS pairwise comparison remained statistically significant after adjustment. Grok showed intermediate performance, whereas DeepSeek showed lower numerical reliability scores, with significant pairwise differences identified only in comparison with ChatGPT and Gemini. Significant differences were observed for mDISCERN [χ² (3)=17.52, p0.05). Readability analysis showed a small but statistically significant difference in FRES among the chatbots, with Gemini demonstrating easier readability than Grok after Bonferroni adjustment, whereas FKGL did not differ significantly among the models; overall, the responses required a high-school to early undergraduate reading level.Conclusion: AI chatbots can generate information with potential clinical relevance in anterior implant dentistry; however, variability in informational reliability persists. Although ChatGPT showed slightly higher numerical scores for some outcomes, the observed differences were generally small, and no chatbot demonstrated clear superiority across all evaluated measures. Expert supervision therefore remains essential before integrating such tools into clinical education.