Machine learning–based prediction of BCG NON-RESPONSE in high-risk NON-muscle-invasive bladder cancer: An exploratory single-centre cohort study
Hakan Şığva, Mehmet Sevim, Kadir Körpe, Vefa Atış, Vedat Arslan, Sadık Görür, Fatih GökalpBackground
Predicting BCG non-response in high-risk non-muscle-invasive bladder cancer (NMIBC) remains a genuine challenge in everyday urological practice. BCG engages the host immune system, and variation in individual immune responses may be reflected in measurable pre-treatment markers.
Objectives
To evaluate exploratory machine learning (ML) models predicting BCG non-response at 12 months from routine pre-treatment data.
Design
Single-centre retrospective observational cohort study.
Methods
We reviewed 114 patients with NMIBC given intravesical BCG at one centre in Turkey, 2019–2024. By the 2021 European Association of Urology classification, all tumours were high grade and all patients high risk. BCG non-response was defined as histologically confirmed recurrence within a fixed 12-month window after at least five induction instillations, with first surveillance cystoscopy at three months. Fourteen candidate predictors, all available before the first instillation, were used; variables determined during or after treatment, including the number of BCG doses, were excluded. Penalized logistic regression, random forest and support vector machine (SVM) models were compared. All preprocessing was fitted within training folds. Performance was assessed by nested cross-validation and reported as pooled out-of-fold area under the curve (AUC) with 95% bootstrap confidence intervals. Reporting follows STROBE and TRIPOD+AI.
Results
Thirty-three patients (28.9%) did not respond; 18 recurrences (54.5%) were high grade. Tumour number (adjusted OR 2.701, 95% CI 1.424–5.122) and tumour size (adjusted OR 1.309 per 10 mm, 95% CI 1.012–1.693) retained an association with non-response. Random forest gave the highest pooled out-of-fold AUC (0.915, 95% CI 0.841–0.974), then SVM (0.874) and penalized logistic regression (0.703). Restricting the outcome to high-grade recurrence reduced random forest discrimination to 0.683–0.754.
Conclusion
ML models discriminated between BCG responders and non-responders at 12 months, but discrimination was substantially lower for high-grade recurrence. Given the modest sample size and the absence of external validation, these findings should be regarded as hypothesis-generating and require prospective multicentre external validation before routine clinical implementation.