American College of Surgeons NSQIP Hospital Benchmarking Using Bayesian Variational Inference to Adjust for Many CPT Codes
Yaoming Liu, Mark E Cohen, Arielle Grieco, Bruce L Hall, Clifford Y KoBackground:
Current ACS NSQIP benchmarking relies primarily on the principal Current Procedural Terminology (CPT) code for procedure-related risk adjustment because conventional regression methods cannot efficiently accommodate numerous sparse CPT variables. Although CatBoost (CATB) can incorporate multiple CPT codes to improve risk estimation, it is computationally intensive and may produce unstable estimates for infrequently performed procedures. This study evaluated Bayesian automatic differentiation variational inference (ADVI) as an alternative approach.
Study Design:
CATB and ADVI were applied to 2019–2023 ACS NSQIP data to estimate CPT-based risk adjustment using all available CPT codes (up to 21 per patient) across 31 outcomes. Performance was compared using calibration (Hosmer-Lemeshow statistic), discrimination (area under the receiver operating characteristic curve), computational time, and effects on hospital-level morbidity benchmarking relative to principal CPT-only risk adjustment.
Results:
Both approaches demonstrated similar discrimination across all 31 outcomes, whereas ADVI consistently achieved superior calibration and markedly lower computational time, with the greatest advantages observed in larger datasets. Among the 15 largest models, mean Hosmer-Lemeshow statistics were 125.23 for ADVI versus 764.72 for CATB, while mean computation times were 9.8 versus 1447.7 minutes, respectively. ADVI and CATB produced similar, although not identical, hospital benchmarking results.
Conclusions:
ADVI is a computationally efficient and statistically robust alternative to CATB for incorporating multiple CPT codes into ACS NSQIP risk adjustment. Improved calibration and more stable estimation for sparse procedure combinations may enhance the reliability and scalability of procedure-based benchmarking.