DOI: 10.3390/bdcc10080274 ISSN: 2504-2289

Benchmarking Supervised Classifiers for Concurrent Multidomain Dropout-Intention Attributions in Higher Education: Evidence from a Colombian Public University

Marieth Agnes Guillen-García, Osnamir Elias Bru-Cordero, Cristian David Correa-Álvarez

Student retention analytics often treats withdrawal as a single outcome, although students may attribute dropout intention to personal, socioeconomic, and academic pressures simultaneously. We benchmarked nine supervised classifiers for identifying a concurrent three-domain attribution profile in a cross-sectional survey of 333 undergraduates at a Colombian public university campus. The response came from a semi-structured weight-allocation item; an audit found that literal label matching altered 21 classifications because of spelling variants and decimal notation. Nine classifiers—logistic regression, decision tree, random forest, neural network, Gaussian Naïve Bayes, k-nearest neighbors, AdaBoost, gradient boosting, and XGBoost—were fitted using six pre-specified predictors. Models were compared by repeated nested stratified cross-validation (five outer folds, three repeats), inner tuning, fold-contained preprocessing, and training-only threshold selection. The concurrent profile occurred in 256 students (76.9%). Logistic regression achieved the highest mean held-out ROC AUC (0.674, 95% CI 0.644–0.704), closely followed by random forest (0.671, 0.641–0.702); their paired difference was nonsignificant after Holm adjustment. Logistic regression had the highest F1 score (0.788), whereas random forest had the highest balanced accuracy (0.617). AdaBoost did not retain its apparent single-holdout advantage. Housing and financial aid had the largest held-out permutation importance. The predictors provided moderate discrimination of a perceptual profile, not a validated prediction of future dropout. Outcome auditing and leakage-free validation materially changed the model ranking.

More from our Archive