DOI: 10.3390/biomedinformatics6040060 ISSN: 2673-7426

Auditing Data Leakage and Temporal Misalignment in Machine-Learning Prognosis of Breast Cancer: A Reproducible Censoring-Aware Reanalysis

Lotfi Tadj, Md Abu Sufian, Wahiba Hamzi, Mai Ali, Amira Ali, Boumediene Hamzi

Background: Near-perfect machine-learning performance on public clinical datasets can arise from data leakage or predictors unavailable at the intended prediction time. We audited a breast cancer vital-status workflow and established a censoring-aware benchmark. Methods: The public SEER-derived dataset contained 4024 women diagnosed in 2006–2010. After removing one duplicate, 616 of 4023 records had status “Dead”. We reproduced a support-vector machine workflow that oversampled the complete dataset before an 80:20 split, quantified exact train-test row overlap, and repeated the analysis with training-only oversampling. A baseline-only support-vector machine and a penalized Weibull accelerated failure-time model were then evaluated using training-only tuning and an untouched test set. Results: Pre-split oversampling placed duplicate feature rows from 50.5% of test observations in training and yielded an area under the receiver operating characteristic curve (AUROC) of 0.998. Splitting first eliminated overlap and reduced AUROC to 0.715 under the retained legacy model. The tuned baseline-only model achieved AUROC 0.745 (95% confidence interval 0.696–0.793). The survival model achieved a test concordance index of 0.753 (0.708–0.797). Conclusions: Resampling order and post-baseline information materially inflated apparent performance. The corrected results support moderate internal prognostic discrimination, not breast cancer diagnosis, triple-negative subtype classification, or clinical deployment.

More from our Archive