Generative Artificial Intelligence in Periodontology: Application of OpenAI Models for the Classification of Periodontitis and Risk Assessment in Patients Undergoing Supportive Periodontal Care
Albert Camlet, Aida Kusiak, Dariusz ŚwietlikBackground/Objectives: Periodontitis is a chronic, inflammatory disease of teeth-supporting tissues and is highly prevalent worldwide. The current classification of periodontal diseases was introduced by the 2017 AAP/EFP World Workshop on the Classification of Periodontal and Peri-Implant Diseases and Conditions. The classification is primarily based on the staging and grading of periodontitis. Periodontal Risk Assessment (PRA), introduced by Lang and Tonetti, is an established tool for estimating the risk of periodontitis progression in patients undergoing supportive periodontal care. Large language models (LLMs) developed by OpenAI are widely investigated for potential clinical applications in medicine and dentistry. The aim of this study was to assess the agreement in PRA and stage and grade classification between clinicians and the following OpenAI models: GPT-5.2, OpenAI-o3, and OpenAI-o4-mini. Methods: The analysis was based on 87 medical records divided into study and control groups (57 and 30 cases, respectively). Records were reviewed for eligibility and diagnostic accuracy by two experienced clinicians. Reference diagnoses were established by consensus. A standardized prompt was used as the instruction for the OpenAI Models. The reference diagnoses were compared with the LLMs’ outputs. Results: LLMs differentiated periodontitis from non-periodontal cases, with sensitivity ranging from 91.2% to 94.7% and a specificity of 100%. Among correctly identified cases of periodontitis, GPT-5.2, OpenAI-o3 and OpenAI-o4-mini achieved high agreement with clinicians in stage and grade assessment (weighted κ values for stage were 0.918, 0.900, and 0.803, respectively; weighted κ values for grade were 0.857, 0.975, and 0.802, respectively), while moderate to substantial agreement was observed for PRA (weighted κ values were 0.622; 0.570; and 0.659, respectively). The analysis including all cases revealed high overall agreement between LLMs and clinicians for stage, grade and PRA, except for GPT-5.2 and OpenAI-o4-mini for grade, which presented substantial and moderate agreement, respectively. Disagreement patterns between LLMs and clinicians indicated that discrepancies occurred mostly between adjacent categories. The global comparison of three LLMs revealed statistically significant differences in agreement with clinicians’ consensus for PRA (χ2 = 24.01; p < 0.001). The paired comparison revealed statistically significant differences between OpenAI-o3 vs. OpenAI-o4-mini for staging (p = 0.006); GPT-5.2 vs. OpenAI-o4-mini for PRA (p = 0.002); and OpenAI-o3 vs. OpenAI-o4-mini for PRA (p < 0.001). Conclusions: OpenAI Models present mostly high agreement with reference diagnoses in periodontitis detection and classification and may potentially support clinical practice. However, their performance in PRA appears to be limited.