Joint Feature Selection and Hyperparameter Optimization Using Evolutionary Algorithms for Diabetes Identification Across Multiple Datasets
Hongli Yan, Junfeng Liao, Shuang GengRoutine clinical data can support the identification of individuals with diabetes, but model performance may change when disease prevalence and feature distributions differ across datasets. We evaluated a machine learning framework that combines clinically informed feature engineering with Differential Evolution (DE)-based joint optimization of feature selection and model hyperparameters. Ten derived features represented nonlinear terms, thresholds, and joint predictive patterns involving glucose, body mass index (BMI), age, and blood pressure. Four DE variants were paired with five classifiers, yielding 20 hybrid models. The Pima Indians Diabetes Dataset (PIDD) was used for development and internal testing, whereas DiaBD and Diabetes_type were used for external evaluation under low (6.5%) and high (78.9%) positive-class prevalences, respectively. MLP+EPSDE achieved the highest values for several performance metrics among the 20 models, such as ACC (0.8117), F1 score (0.7339), AUC (0.8318), and MCC (0.5889). Its external AUCs were 0.7211 on DiaBD and 0.8964 on Diabetes_type, indicating that ranking discrimination was retained across markedly different class distributions. Because accuracy is prevalence-dependent, the external results were not interpreted as evidence of probability calibration or clinical utility. SHapley Additive exPlanations (SHAP) identified Glucose_Squared, BMI_Category, Age_Group, and Age as the most influential variables in MLP+EPSDE; these findings describe model-based predictive contributions rather than causal or biological interactions. Overall, the framework provides an interpretable approach to diabetes identification in developing countries, although prospective validation is required before clinical application.