Development and Clinical Validation of an Automated LLM Judge for Evaluating Perioperative Patient Questions
Bernardo G. Collaco, Nadia G. Wood, Yunguo Yu, Carina Rosa Malena, Srinivasagam Prabha, Ashton L. Boon, Zhihui Fang, Anjali Bhagra, Bradley C. Leibovich, Antonio Jorge ForteBackground: Large language models (LLMs) are increasingly used in patient-facing clinical applications, creating a need for scalable methods to evaluate the accuracy, safety, and appropriateness of their responses. Although physician review remains the reference standard, it is resource-intensive and difficult to scale. Objective: To develop and clinically validate an LLM-as-a-judge framework for evaluating responses to patients’ periprocedural questions. Methods: We retrospectively evaluated 1285 patient–AIVA interactions from 112 patients across two Mayo Clinic sites and multiple surgical specialties. Physician reviewers classified interactions using a predefined true-positive (TP), false-negative (FN), true-negative (TN), and false-positive (FP) framework. The finalized LLM judge independently evaluated the same interactions. Agreement was assessed using four-class and category-specific agreement, Cohen’s kappa, and a secondary binary analysis of response correctness. Results: Overall, four-class agreement was 93.1% (95% CI, 91.7–94.5%), with Cohen’s κ = 0.852 (95% CI, 0.823–0.882). Category-specific agreement was 94.1% for TP, 90.7% for FN, 95.5% for TN, and 84.9% for FP classifications. In the binary correctness analysis, accuracy was 93.4% (95% CI, 92.0–94.7%), sensitivity 94.6%, specificity 90.2%, precision 96.2%, and F1-score 0.954. McNemar’s test demonstrated no significant asymmetry between paired classifications (p = 0.11). Conclusions: A clinically grounded LLM judge closely approximated physician evaluation of patient-facing AI responses. These findings support its potential as a scalable assistive monitoring tool while preserving physician oversight for uncertain or safety-sensitive interactions.