DOI: 10.3390/jcm15197397 ISSN: 2077-0383

Comparative Evaluation of Large Language Model Interfaces in Third Molar Surgery Complication Scenarios: Response Quality, Clinical Content, Potential Clinical Risk, Readability, and Externally Observable Response Latency—A Cross-Sectional Comparativ

İnci Rana Karaca, Esmanur Başer, Emre Ulubaş, Elif Betül Yıldırım

Background/Objectives: Large language models (LLMs) are increasingly investigated in dentistry and oral and maxillofacial surgery, but comparative evidence regarding specialist-level response quality, clinical content, and potential safety concerns remains limited. This study compared four contemporary LLM interfaces in third molar surgery complication scenarios. Methods: Thirty open-ended specialist-level scenarios were submitted once to ChatGPT 5.4 Thinking, DeepSeek-V3.2, Gemini 3.1 Pro, and Microsoft Copilot (Work and Learning mode) in isolated single-turn sessions, yielding 120 responses. Two oral and maxillofacial surgeons, blinded to interface identity, independently assessed responses using a task-adapted five-point Global Quality Score (GQS). An additional post hoc criterion-referenced evaluation assessed diagnosis, management, escalation/referral, potentially harmful recommendations, safety-critical omissions, and potential clinical risk. Readability and externally observable response latency were also evaluated. Results: Among the sampled outputs, mean GQSs were 4.53 for Gemini, 4.40 for ChatGPT, 4.30 for DeepSeek, and 3.92 for Copilot (p < 0.001). Mean Clinical Content Scores were 4.80/5 for ChatGPT, 4.77 for Gemini, 4.75 for DeepSeek, and 3.90 for Copilot (p < 0.001). All responses were rated as recognizing the target complication. Within the 120 sampled outputs, one Copilot response met the study-specific criterion for a potentially harmful recommendation, and two Copilot responses met the predefined criteria for safety-critical omissions. Within this benchmark, the sampled Copilot outputs had higher study-specific potential clinical-risk scores than the sampled outputs from the other evaluated interfaces (p < 0.001). Readability and externally observable response latency also differed across the sampled interface outputs. Conclusions: In this single-generation benchmark, the sampled outputs showed different profiles in specialist-rated quality, clinical content, potential clinical risk, readability, and externally observable response latency under the specific interface configurations tested. These findings apply to the observed responses from the single synchronized testing session and should not be interpreted as establishing superiority of any underlying LLM, independent diagnostic accuracy, clinical equivalence, clinical safety, effectiveness, or readiness for autonomous decision support.