DOI: 10.1097/xcs.0000000000002238 ISSN: 1072-7515

Matching Surgical Datasets to Prediction Tasks: NSQIP and Non-NSQIP Artificial Intelligence Models Across Specialties

Divya Kewalramani, Justin Benton, Zack Barta, Lily Wushanley, Amanda Teichman, Joseph Hanna, Amin Madani, Philip S Barie, Tyler Loftus, Mayur Narayan

Background:

Surgical artificial intelligence models depend on training data that may not capture predictors required for specific outcomes. We compared discrimination of models derived from the American College of Surgeons National Surgical Quality Improvement Program (NSQIP) with models using non-NSQIP datasets across surgical specialties and prediction tasks.

Study Design:

Structured searches of PubMed, Scopus, and EMBASE identified adult surgical prediction models published from October 1, 2015, through January 31, 2026. The primary metric was area under the receiver-operating characteristic curve (AUROC). Multivariable ordinary least squares regression with study-level cluster-robust standard errors estimated adjusted differences in AUROC between data sources.

Results:

Seventy-five studies contributed 295 models: 145 NSQIP and 150 non-NSQIP. Mean AUROC was 0.71±0.09 for NSQIP vs 0.80±0.10 for non-NSQIP models (p<0.001). After adjustment for publication year, outcome, validation type, variable count, specialty, architecture, log-transformed sample size, and data modality, NSQIP remained associated with lower AUROC (β=−0.056, 95% CI −0.105 to −0.007; p=0.026). The difference was largest for infection prediction (median AUROC 0.68 [IQR 0.64-0.70] vs 0.98 [0.86-0.99]; p<0.001) and absent for mortality (0.79 [0.68–0.92] vs 0.80 [0.72-0.85]; p=0.94). Precision weighting attenuated the adjusted association (β=+0.008, 95% CI −0.093 to 0.109; p=0.876).

Conclusions:

Discrimination differences between NSQIP- and non-NSQIP-derived surgical prediction models were outcome-dependent and sensitive to precision weighting. These findings support matching training datasets to prediction tasks rather than assuming database size or standardization ensures superior discrimination.