DOI: 10.3390/children13081043 ISSN: 2227-9067

Machine Learning-Based Prediction of Six-Minute Walk Distance in Children with Obesity

Emna Makni, Mohamed Elloumi, Achraf Ammar, Mehdi Ben Brahim, Younes Hachana

Background: The six-minute walk test (6MWT) assesses functional exercise capacity, but existing reference equations for children with obesity rely on traditional linear regression, potentially overlooking complex, non-linear relationships between anthropometric characteristics and functional exercise capacity. Objective: This study aimed to develop and internally validate machine-learning (ML) prediction models and preliminary prediction equations derived from explainable ML models for six-minute walk distance (6MWD) in Tunisian school-aged children with obesity and to compare their predictive performance with a conventional regression-based approach. Methods: We analyzed data from 236 school-aged children with obesity (104 females, 132 males; 6–12 years). Anthropometric measurements included body mass (BM), height, body mass index (BMI), waist circumference (WC), and hip circumference (HC). Five models were evaluated: linear regression, Ridge, Lasso, Elastic Net, and XGBoost. Performance was assessed using five-fold cross-validation and evaluated by the mean absolute error (MAE) and root mean square error (RMSE). Predictor importance was assessed using SHapley Additive exPlanation (SHAP) and Gini importance. Results: XGBoost achieved the best predictive performance, with the lowest MAE (17.4 ± 2.5 m in females and 19.5 ± 1.9 m in males) and RMSE (25.3 ± 3.1 m in females and 26.4 ± 2.2 m in males). Age was the strongest predictor across all models (SHAP: 54.2–62.0%; Gini importance: 0.52–0.69), followed by height and BMI. Sex-specific analyses indicated that, in females, age and BMI contributed ~80% to the cumulative SHAP analysis; whereas, in males, age, height, and WC were the primary factors. Conclusions: ML, particularly XGBoost, significantly improves 6MWD prediction in school-aged children with obesity compared with traditional linear regression. Explainable ML increases model interpretability by evaluating the relative importance of anthropometric predictors. These obesity-specific prediction models may better capture complex non-linear associations between anthropometrics and functional exercise capacity. These initial population-specific models should be validated in larger independent samples before routine use clinical or field settings.

More from our Archive