DOI: 10.1177/20552076261492994 ISSN: 2055-2076

Assessing the accuracy of AI-generated responses to periodontal scaling and root planing patient FAQs

Ahmed M. Kabli

Background/Objective

Periodontal diseases affect nearly half the global population, making patient education critical for successful outcomes. With increasing public reliance on artificial intelligence (AI) for health information, evaluating the accuracy of AI-generated responses to periodontal questions is essential. To assess the clinical accuracy, reliability, and consistency of ChatGPT and Gemini in answering frequently asked questions (FAQs) regarding scaling and root planing (SRP).

Methods

Fifty binary (yes/no) FAQs about SRP were developed and posed to ChatGPT and Gemini across three independent sessions (morning, afternoon, night), yielding 300 total responses. Two board-certified periodontists evaluated responses against current AAP/ADA guidelines. Accuracy, consistency, domain-specific performance, and inter-model agreement were analyzed using descriptive statistics, an exact two-sided McNemar test, chi-square, Cochran’s Q , and Cohen’s Kappa.

Results

Gemini demonstrated significantly higher accuracy than ChatGPT (99.3% vs. 93.3%; exact two-sided McNemar test, p = 0.012). Domain-specific accuracy ranged from 86.7% to 100% for ChatGPT and 96.3% to 100% for Gemini. Consistency rates were 92.0% for ChatGPT and 98.0% for Gemini ( p = 0.36). ChatGPT showed significant accuracy improvement across sessions (90.0% to 98.0%, p = 0.04). Inter-model agreement was 92.7% with Kappa = -0.012 ( p = 0.78).

Conclusions

Both AI models demonstrated high clinical accuracy, with Gemini significantly outperforming ChatGPT. Consistency was good for both models. AI tools may serve as useful supplementary tools for delivering basic information. However, AI platforms must complement, and never replace, expert clinical judgment, professional diagnosis, or individualized patient consultations.