Evaluation of Small and Large Language Models for Calculation of the ASA Score and Charlson Comorbidity Index in Orthopedic Surgical Patients: A Retrospective Concordance Analysis
Marco Di Maio, Giorgio Stopper, Vincenzo Di Matteo, Katia Chiappetta, Guido Grappiolo, Mattia LoppiniBackground: The ASA Physical Status (ASA-PS) classification and the Charlson Comorbidity Index (CCI) are common pre-operative scoring tools. Language models could automate structured pre-operative scoring, but direct comparisons require paired inference because all models are evaluated on the same patients. Methods: In this retrospective single-center concordance analysis, 101 consecutive adult orthopedic patients were independently rated by two clinicians; the rounded mean for ASA-PS and arithmetic mean for CCI formed a clinician-derived composite reference. The cohort contained no ASA-PS IV-V patients. Six model configurations received identical prompts. Agreement was assessed using quadratic weighted kappa, ICC(2,1), exact and adjacent agreement, MAD, RMSE, and Bland–Altman limits. Post hoc between-model comparisons used 10,000 patient-level paired bootstrap replicates with Benjamini–Hochberg correction. Results: Inter-clinician weighted kappa was 0.713 for ASA-PS and 0.914 for CCI. GPT-5.2 reached kappa 0.884 for ASA-PS and 0.970 for CCI. In paired analyses, GPT-5.2 had significantly higher quadratic weighted kappa than every other tested model for both outcomes and significantly higher ICC for CCI after multiplicity correction. Phi4 and deepseek-r1-70B were not significantly different from inter-clinician agreement for CCI kappa or ICC; equivalence was not tested. Conclusions: Among the six evaluated model configurations, GPT-5.2 achieved significantly higher agreement with the clinician-derived composite reference than the other tested models for both ASA-PS and CCI in post hoc paired analyses with multiplicity correction. Locally deployable phi4 and deepseek-r1-70B showed CCI agreement estimates that were not statistically distinguishable from inter-clinician agreement, although equivalence was not tested. These findings are limited to the evaluated models and study cohort.