Machine Learning-Based Imputation for Breast Cancer Prediction: Evaluating Performance Under Complex Missing Data Mechanisms
Nyatuga Gideon Nyakundi, John Ndiritu, Ivivi Joseph Mwaniki, Timothy Kevin KamanuMissing data remain a major challenge in breast cancer research because they can introduce bias, reduce statistical efficiency, and compromise the performance of predictive models. Although numerous imputation techniques have been proposed, their comparative performance under different missing-data mechanisms and their impact on downstream classification remain inadequately understood. This study systematically compared statistical and machine learning-based imputation methods using two publicly available breast cancer datasets representing complementary clinical settings. The methods were evaluated under simulated Missing Completely at Random (MCAR), Missing at Random (MAR), and Missing Not at Random (MNAR) mechanisms using both reconstruction accuracy and downstream classification performance. The results showed that no single imputation method consistently achieved the best performance across both datasets. Regularized regression and machine learning-based methods generally outperformed conventional statistical approaches, although the optimal method depended on the characteristics of the dataset. Furthermore, the best-performing imputation methods preserved downstream classification performance despite the introduction of missing data, demonstrating that reconstruction accuracy alone is insufficient for selecting imputation strategies intended for predictive modelling. Overall, the findings highlight the importance of considering dataset characteristics, missing-data mechanisms, and the intended analytical objective when selecting imputation methods. The proposed evaluation framework provides a robust approach for assessing missing-data handling strategies in breast cancer prediction studies and other biomedical machine learning applications.