DOI: 10.32322/jhsm.1914765 ISSN: 2636-8579

A comparative evaluation of Artificial Intelligence Chatbots’ responses in special needs dentistry

Hamit Tunç
Aims: This pilot study aimed to evaluate the accuracy and consistency of widely used Artificial Intelligence (AI) chatbots in responding to questions and diagnosing syndromes related to special needs dentistry. Methods: Three AI chatbots (ScholarGPT, ChatGPT, and Gemini) were assessed using ten true/false questions and ten diagnostic case-based questions derived from scientific literature. Each chatbot was queried three times over a three-week period. Responses were independently evaluated by two pediatric dentistry residents. Accuracy was scored dichotomously, while inter-reviewer reliability was measured using intraclass correlation coefficients (ICC). Statistical comparisons were conducted using the Wilcoxon exact test, and internal consistency was assessed with Cronbach’s alpha. Results: All chatbots demonstrated high inter-reviewer agreement (ICC >0.85, p