DOI: 10.4258/hir.2026.32.3.232 ISSN: 2093-369X

Application of Machine Learning Algorithms for Predicting Infant Mortality in India: An Analysis of the National Family Health Survey-5, 2019–2021

Jaya Prasad Tripathy

Objectives: Machine learning (ML) techniques have shown strong potential for predicting infant mortality (IM), but their application in the Indian context remains limited. This study aimed to use ML algorithms to predict IM in India using a large national survey database.Methods: Data were analyzed from the National Family Health Survey-5, 2019–2021, a large cross-sectional survey. Random forest, decision tree, adaptive boosting, logistic regression, and naïve Bayes models were implemented using Weka version 3.8.3. Model performance was evaluated using accuracy, precision, F1-score, Matthews correlation coefficient, and area under the curve (AUC).Results: Compared with logistic regression, which achieved 65.2% accuracy, the random forest and decision tree models showed higher predictive accuracy, at 74.1% and 73.2%, respectively. Their AUCs were 0.80 and 0.79, respectively, compared with 0.69 for logistic regression. The models identified birth order, maternal education, twin birth, wealth index, cooking fuel use, and age at first birth as the six strongest predictors of IM.Conclusions: In this analysis, random forest and decision tree models outperformed logistic regression in predicting IM. These findings underscore the relevance of sociodemographic and economic disparities and support the use of ML algorithms for risk prediction and targeted interventions aimed at reducing IM.

More from our Archive