DOI: 10.1111/edt.70113 ISSN: 1600-4469

Accuracy of Chatbot‐Based Health Advice Related to Traumatic Dental Injuries: A Systematic Review and Meta‐Analysis

Nitesh Tewari, Amna Afrin Kidwai, Ekta Wadhwani, Mugilan Ravi, Georgios Tsilingaridis, Priscilla Barbosa Ferreira Soares, Paulo César Bandeira Junqueira, Caroline Garcia Orsi, Partha Haldar, Carlos José Soares

ABSTRACT

Background/Aims

The growing use of generative artificial intelligence, especially chatbots, has motivated researchers to test the accuracy of health advice provided. Though there are studies comparing chatbot health advice for traumatic dental injuries (TDI), variable results have been reported. Hence, this systematic review aimed to assess contemporary chatbots for their ability to accurately and reliably provide responses to different types of questions related to emergency care of TDI.

Methods

The protocol was developed using the Cochrane Handbook and the Chatbot‐Assessment‐Reporting‐Tool (CHART) statement and registered in the Open Science Framework. A literature search, based on the research question, was conducted in PubMed, EMBASE, Scopus, and Web of Science on September 19th, 2025, without any limitations. Screening of titles and abstracts and later full text, and data extraction were performed. These steps were followed by qualitative synthesis and meta‐analysis. The quality of studies was assessed using their compliance with CHART.

Results

The review included 10 articles published in 2024 and 2025. Seven contemporary chatbots were evaluated with different themes and types of questions. Eight of the studies showed good compliance with CHART (≥ 75%). The overall pooled accuracy was 65% (95% CI 52%–78%, I 2  = 88%) in ChatGPT‐3.5, 79% (95% CI 61%–93%, I 2  = 88.8%) in ChatGPT‐4.0, and 70% (95% CI 60%–79%, I 2  = 86.4%) in Gemini.

Conclusion

Overall accuracy of chatbot‐generated advice was moderate, with substantial variation depending on the type of questions assessed. The meta‐analysis demonstrated a suggestive trend of higher accuracy for ChatGPT‐4.0, followed by Gemini, Copilot P, and ChatGPT‐3.5. However, these findings should not be interpreted as evidence of clinical reliability, as even limited inaccuracies in emergency dental trauma advice may contribute to delayed or inappropriate management and adverse outcomes.

More from our Archive