DOI: 10.4103/jips.jips_182_26 ISSN: 0972-4052

Interobserver reliability and temporal variability of artificial intelligence chatbots in Kennedy-classified removable partial denture design: A cross-sectional analysis

Shady M. El Naggar, Ahmed M. Esmat, Eman G. Helal, Maie F. Khalil, Ayman M. Gouda

Abstract

Aims:

This study aimed to investigate the consistency of ratings, the quality of chatbot-generated removable partial denture (RPD) designs, and the repeatability of chatbot outputs over time through standardized Kennedy classification-based RPD design prompts.

Settings and Design:

This study used a cross-sectional design to test the consistency of four AI chatbots among raters (ChatGPT; GPT-4; OpenAI, Copilot 1.25062.106.0; Microsoft, AI Chat-AI language model; DeepAI, Claude AI; Claude 3 Opus; Anthropic) on RPD design for 80 virtual Kennedy-classified cases (40 maxillary/40 mandibular).

Materials and Methods:

Eighty partially edentulous virtual cases (Class I, I Mod 2, II, II Mod 1, II Mod 2, III Mod 1, III Mod 2, and IV) were assigned to AI chatbots. Each case was randomly assessed three times daily, resulting in 960 responses evaluated. Four removable prosthodontists scored responses using a modified global quality score based on a five-point Likert scale.

Statistical Analysis Used:

Statistical analyses used the intraclass correlation coefficient (ICC) and two-way Analysis of variance ANOVA ( P ≤ 0.05). Interobserver reliability was measured using the ICC, technical quality was compared with an expert consensus reference standard, and repeated-measures ANOVA demonstrated significant temporal variation in chatbot outputs.

Results:

All chatbots showed excellent interobserver reliability (ICC >0.90), but the resulting RPD designs had lower scores compared with the expert gold reference standards. Repeated-measures analysis indicated significant intra-rater variability across all chatbots ( P < 0.001), suggesting lower repeatability.

Conclusions:

AI chatbots showed excellent interobserver reliability but limited technical quality and temporal stability in the RPD framework design. ChatGPT (GPT-4; OpenAI) performed best, but all models still need expert supervision and further optimisation before use.