DOI: 10.4258/hir.2026.32.3.214 ISSN: 2093-369X

Evaluation and Comparison of Machine Learning Methods for Type 2 Diabetes Classification and Associated Factors

Masoumeh Dadashpour, Mohammad Yousefi, Sohrab Effati, Somayeh Ghiasi Hafezi

Objectives: Type 2 diabetes mellitus (T2DM) is a prevalent chronic metabolic disorder associated with serious complications, including nephropathy, cardiovascular disease, retinopathy, and neuropathy. Given its increasing incidence and the complexity of associated factors—such as obesity, metabolic syndrome, and sedentary lifestyle—accurate identification is essential. This study aimed to evaluate and compare the performance of several machine learning algorithms to identify key associated factors and detect individuals with T2DM within this dataset.Methods: A publicly available dataset from Kaggle, comprising health records of 99,982 individuals, was used. Five supervised machine learning models were evaluated: Bayesian ridge regression, logistic regression, extreme gradient boosting (XGBoost), artificial neural networks, and random forest. Each model was trained and evaluated to assess classification performance. Performance was measured using the area under the receiver operating characteristic curve (AUC–ROC) and accuracy. SHapley Additive Explanations (SHAP) values were used to interpret model outputs and identify the most influential features.Results: Among the five models, XGBoost demonstrated the highest performance, achieving an accuracy of 96% and an AUC–ROC of 0.98. SHAP analysis identified hemoglobin A1c, blood glucose, age, body mass index, and sex as the most influential predictors of T2DM. Conclusion: XGBoost was the most effective algorithm for identifying individuals with T2DM in this dataset. It also provided insights into the relative importance of clinical features, supporting more precise classification. However, results should be interpreted with caution until validated in independent cohorts.

More from our Archive