Artificial intelligence-driven prediction of neoadjuvant chemotherapy response in adult patients with gastric cancer: a systematic review and meta-analysis
Yuntian Deng, Anran Bao, Hongshan WangObjectives
A substantial proportion of adults with locally advanced gastric cancer derive limited benefit from neoadjuvant chemotherapy (NAC), with a non-response rate of 30%–50%. Artificial intelligence (AI) models that integrate imaging, pathology and clinical data seek to predict pretreatment NAC response to guide decision-making. We conducted a systematic review and meta-analysis to evaluate the predictive performance of AI models for NAC response in adults with gastric cancer.
Design
Systematic review and meta-analysis in accordance with Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines.
Data sources
PubMed, Embase, Scopus, Web of Science and ScienceDirect were searched from inception to 30 June 2025, without language restrictions.
Eligibility criteria
Studies that developed or internally/externally validated AI models (deep learning or non-linear machine learning) to predict NAC response in adults (≥18 years) with gastric cancer.
Data extraction and synthesis
Data extraction followed the CHARMS (CHecklist for critical Appraisal and data extraction for systematic Reviews of prediction Modelling Studies) checklist. Model quality and risk of bias were appraised with the Prediction Model Study Risk of Bias Assessment Tool for AI (PROBAST+AI). Where two or more independent validation datasets reported outcomes—preferably mappable to major pathological response (tumour regression grading 0–1/Becker 1a–1b)—we pooled area under the receiver operating characteristic curve (AUC) using random-effects models and quantified heterogeneity with I². Certainty of evidence was graded with the Grading of Recommendations Assessment, Development and Evaluation (GRADE).
Results
17 studies met inclusion criteria; 16 (94%) were conducted in China and nine were multicentre. 14 cohorts contributed to internal-validation meta-analyses and 10 to external-validation meta-analyses. Pooled AUC was 0.844 (95% CI 0.812 to 0.877, I²=17.3%) for internal validation and 0.812 (95% CI 0.775 to 0.848, I²=52.9%) for external validation. Subgroup analysis indicated that data type was the strongest modifier of pooled area under the curve in external validation (p=0.008, R²=99.8%). Only two studies directly compared AI models against traditional clinical prediction models with AUC (0.823, 95% CI 0.782 to 0.863 vs 0.598, 95% CI 0.521 to 0.675). By PROBAST+AI, all 17 studies raised high or unclear concerns regarding development quality (13 high, 4 unclear), and risk of bias in evaluation was high in 12 studies, low in 2, and unclear in 3. GRADE-rated certainty of evidence was low for internal validation and very low for external validation.
Conclusions
Across 17 studies, AI models demonstrated moderate to good discrimination for predicting NAC response in gastric cancer, with pooled AUCs of 0.844 (95% CI 0.812 to 0.877) for internal validation and 0.812 (95% CI 0.775 to 0.848) for external validation. Model-methodology subgroup analyses did not demonstrate robust differences between deep learning and machine-learning approaches, and AI models significantly outperformed traditional clinical assessment. However, GRADE-rated certainty of evidence was low for internal validation and very low for external validation, reflecting high risk of bias on PROBAST+AI, limited external validation, and geographic concentration of studies (16/17 from China). Pending prospective, multicentre evaluation with harmonised endpoints and routine calibration and decision-curve reporting, these models should serve as decision-support tools rather than standalone arbiters of care.
PROSPERO registration number
CRD420251025470.