Non-Invasive Sociodemographic Cues for Type 2 Diabetes Risk Stratification in South Africa
Reza Fahimi, Abdulaziz Alrubayyi, Habibullah Muhammad Kamal, Ibrahim Khalil Ja’afar, Mohanad Alhothify, Thomas M. Barber, Olalekan A. UthmanBackground/Objectives: Type 2 diabetes mellitus (T2DM) is a growing burden in low- and middle-income countries, and non-invasive risk tools have been proposed where laboratory screening is not feasible. Fast-and-frugal decision trees (FFDTs) are transparent, non-compensatory rules that make sensitivity–specificity trade-offs explicit. We asked how far five routinely collected sociodemographic cues can go for classifying self-reported diabetes status and benchmarked FFDTs against the simplest possible alternatives. Methods: A cross-sectional secondary analysis of the 2016 South Africa Demographic and Health Survey Adult Health recode (N = 10,292; 459 self-reported cases, 4.46%) was carried out. Five binary cues were derived: age category, household wealth, employment, education and sex. Data were split 70/30 into development and a locked hold-out set. Three purpose-specific FFDTs were selected by 5 × 10-fold cross-validation within the development data only, frozen, and then evaluated once on the hold-out set. Comparators were a pre-specified age-only rule and logistic regression using the same five predictors at a development-derived Youden threshold. We report survey-weighted estimates, internal–external cross-validation across nine provinces, external validation in 12 further DHS, subgroup performance, and a log–log analysis of error against training sample size. Results: On the hold-out set the balanced screening FFDT achieved sensitivity 0.920 (95% CI 0.862–0.959), specificity 0.585 (0.567–0.602) and balanced accuracy 0.752 (0.727–0.776). The pre-specified age-only rule achieved balanced accuracy 0.751 (0.724–0.775), while logistic regression at the development-derived Youden threshold achieved 0.766 (0.735–0.797). Across 50 independent splits, the case-finding FFDT was arithmetically identical to the age-only rule in every split; the balanced FFDT exceeded it by a median of 0.001 (95% range −0.004 to +0.025); the referral-minimisation FFDT was worse by a median of 0.124. Performance was stable across provinces (balanced accuracy 0.731–0.801) but degraded across countries (0.508–0.724), with referral rate ranging from 0.225 to 0.806. Error scaled with training size as beta = −0.044 (95% CI −0.063 to −0.025). Conclusions: Transparent decision trees built from non-invasive sociodemographic cues classify self-reported diabetes status no better than a single question about age, do not transport reliably across countries, and are unlikely to be materially improved by simply increasing the training sample size. Their value lies in auditability rather than accuracy. Progress in low-burden risk stratification for T2DM in these settings will require cues that carry a biological rather than health system signal.