Patient-Facing AI Chatbot Treatment-Direction Advice in Orthodontic Health Communication: A Scenario-Based Comparison with Expert Consensus
Neslihan Karaoğlan, Hakan KaraoğlanBackground/Objectives: AI chatbots may shape patient expectations before professional consultation. This scenario-based first-response study evaluated whether four user-facing chatbots provided orthodontic treatment-direction advice concordant with an expert benchmark and whether responses contained safety, referral, or overconfidence concerns. Methods: Forty fictional Turkish patient-oriented scenarios across eight categories were independently coded by three orthodontists as clear aligners, fixed appliances, both options, examination required, or advanced specialist/surgical evaluation required. Each scenario was submitted once to ChatGPT, Claude, Copilot, and Gemini on 20 May 2026. Two independent non-author orthodontists coded 160 archived first responses using a predefined framework, with adjudication before analysis. Results: Inter-expert agreement was moderate (Fleiss kappa = 0.491; Gwet AC1 = 0.528). Under the majority benchmark, exact concordance was 82.5% for ChatGPT, 67.5% for Claude, 42.5% for Copilot, and 37.5% for Gemini (Cochran Q = 34.105, p < 0.001). The overall difference remained significant in the 17 unanimous scenarios (Q = 11.455, p = 0.010), but a post hoc alternative-reference analysis that adopted the dissenting expert code in the 23 non-unanimous scenarios attenuated the rates to 57.5%, 52.5%, 52.5%, and 42.5%, respectively (Q = 4.222, p = 0.238). Coded safety-concern rates ranged from 15.0% to 62.5%. Conclusions: The sampled first responses differed in treatment direction and safety coding, but estimates were sensitive to the expert reference definition. Under the tested single-date, single-language, and single-run conditions, the findings represent a conditional snapshot rather than a time-invariant ranking of model capability. Patient-facing chatbots should support nondirective pre-consultation education and referral, not autonomous appliance selection.