Interobserver reliability and temporal variability of artificial intelligence chatbots in Kennedy-classified removable partial denture design: A cross-sectional analysis
Shady M. El Naggar, Ahmed M. Esmat, Eman G. Helal, Maie F. Khalil, Ayman M. GoudaAbstract
Aims:
This study aimed to investigate the consistency of ratings, the quality of chatbot-generated removable partial denture (RPD) designs, and the repeatability of chatbot outputs over time through standardized Kennedy classification-based RPD design prompts.
Settings and Design:
This study used a cross-sectional design to test the consistency of four AI chatbots among raters (ChatGPT; GPT-4; OpenAI, Copilot 1.25062.106.0; Microsoft, AI Chat-AI language model; DeepAI, Claude AI; Claude 3 Opus; Anthropic) on RPD design for 80 virtual Kennedy-classified cases (40 maxillary/40 mandibular).
Materials and Methods:
Eighty partially edentulous virtual cases (Class I, I Mod 2, II, II Mod 1, II Mod 2, III Mod 1, III Mod 2, and IV) were assigned to AI chatbots. Each case was randomly assessed three times daily, resulting in 960 responses evaluated. Four removable prosthodontists scored responses using a modified global quality score based on a five-point Likert scale.
Statistical Analysis Used:
Statistical analyses used the intraclass correlation coefficient (ICC) and two-way Analysis of variance ANOVA (
Results:
All chatbots showed excellent interobserver reliability (ICC >0.90), but the resulting RPD designs had lower scores compared with the expert gold reference standards. Repeated-measures analysis indicated significant intra-rater variability across all chatbots (
Conclusions:
AI chatbots showed excellent interobserver reliability but limited technical quality and temporal stability in the RPD framework design. ChatGPT (GPT-4; OpenAI) performed best, but all models still need expert supervision and further optimisation before use.