An Electronic Health Record–Integrated, Large Language Model–Powered Tool to Triage Surgical Patients
Jane Wang, Timothy Keyes, April S. Liang, Stephen P. Ma, Jason Shen, Jerry Liu, Nerissa Ambers, Abby Pandya, Rita Pandya, Jason Hom, Natasha Steele, Jonathan H. Chen, Kevin SchulmanImportance
Surgical comanagement (SCM) is an evidence-based care model in which hospitalists jointly manage medically complex perioperative patients alongside surgical teams. Despite its clinical and financial value, effective use of SCM is limited by the need to manually identify eligible patients; large language models (LLMs) are increasingly prevalent in clinical workflows and could be useful in selecting patients for SCM.
Objective
To assess whether SCM eligibility triage can be automated.
Design, Setting, and Participants
This prospective, unblinded quality improvement study was conducted at Stanford Health Care from September 2025 to February 2026. An LLM-based, electronic health record (EHR)–integrated, human-in-the-loop surgical triage tool (SCM Navigator) provided SCM triage recommendations, followed by physician review. All SCM attending physicians were invited to participate.
Exposure
Using preoperative documentation, structured data, and clinical criteria for perioperative morbidity, the tool categorized patients as appropriate, not appropriate, or possibly appropriate for SCM. SCM faculty indicated clinical judgment of SCM appropriateness and provided free-text feedback when they disagreed with the tool.
Main Outcomes and Measures
The sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) of the tool were measured using physician feedback as a reference. Free-text reasons were thematically categorized. Manual medical record review was conducted on all false-negative cases. For the largest false-positive category, manual medical record reviews were conducted on 15 randomly selected cases subsequently seen by SCM and 15 randomly selected cases not seen by SCM.
Results
Overall, 14 of 16 SCM attending physicians participated in the study. Among 6193 triaged surgical cases (median [IQR] age, 60.2 [41.6-71.3] years; 3036 [49.0% female]), 1582 (25.5%) were recommended for hospitalist consultation. Using treating physicians’ determinations as the reference standard, the tool had high sensitivity (0.94; 95% CI, 0.91-0.96) and moderate specificity (0.74; 95% CI, 0.71-0.77) for identifying patients appropriate for SCM. Post hoc medical record review suggested that most discrepancies reflected modifiable gaps in clinical criteria, institutional workflow, or physician practice variability rather than LLM misclassification, which accounted for 2 of 19 false-negative cases (11%).
Conclusions and Relevance
In this prospective quality improvement study of an LLM-powered, EHR-integrated, human-in-the-loop AI system, AI-enabled screening tools accurately augmented surgical patient triage.