Leakage-Controlled Machine Learning for Territorial Epidemiological Risk Stratification: A Biomedical Informatics Framework Using Administrative Health Data in the Colombian Orinoquía
Roberto Ferro Escobar, Danilo Alberto Vera ParraDigital health observatories require predictive frameworks that transform administrative health data into reliable territorial risk indicators. We developed a biomedical informatics framework for machine-learning-based epidemiological risk stratification in the Colombian Orinoquía, with explicit attention to data-leakage prevention, validated population denominators, and unit-of-analysis discipline. From 354,088 morbidity records (2018–2023; Arauca, Casanare, Meta, Vichada) we derived 20,212 independent strata (municipality × year × diagnostic group × sex × age category × health component) and computed morbidity rates using official population projections from the Colombian National Administrative Department of Statistics (DANE), correcting a denominator instability present in the original extract. After excluding leakage-generating variables, Gradient Boosting, Random Forest, and a one-hot Logistic Regression baseline were evaluated through temporal validation, group-aware cross-validation, leave-one-department-out validation, and a geographic ablation experiment. Under leakage-controlled, stratum-level conditions, Gradient Boosting achieved an area under the receiver operating characteristic curve (AUC-ROC) of 0.913 [95% confidence interval (CI): 0.903–0.923] (Brier: 0.106). SHapley Additive exPlanations (SHAP) analysis identified diagnostic group and age category as the dominant predictors, supported by broadly stable performance without geographic identifiers (AUC 0.871) and consistent cross-department transferability (0.807–0.889); a residual municipality-level clustering effect is reported transparently. Validated denominators reversed the apparent territorial gradient, with the most remote department exhibiting the lowest documented morbidity, consistent with under-registration.