Cardiovascular Classification in Middle‐Aged and Elderly
CKD
Patients: A Machine Learning Approach in China and the United States
Yan Zhang, Kaihang Xu, Haoyao Zhang, Yujian Pan, Tiansheng Zhu ABSTRACT
Background
Cardiovascular disease (CVD) drives mortality in chronic kidney disease (CKD) patients. Conventional risk tools underperform in CKD, and few machine learning (ML) models undergo cross‐national validation. We developed interpretable ML models for classifying CVD amongst Chinese and US middle‐aged and elderly CKD patients.
Methods
We conducted retrospective cross‐sectional analyses using two nationally representative cohorts: NHANES (2011–2016, US) and CHARLS (Wave 1 and 3, China). A total of 1057 CKD patients from NHANES and 1109 from CHARLS were included. Seven ML algorithms were trained and optimized. Model performance was assessed on internal test sets and via bidirectional cross‐population external validation. Model interpretability was achieved using SHAP (SHapley Additive exPlanations) analysis.
Results
The XGBoost model performed best in NHANES (test AUC: 0.753), while the Logistic Regression model was optimal for CHARLS (test AUC: 0.769). During bidirectional external validation, both models showed a decline in discriminative performance when deployed across populations, indicating limited cross‐population transportability. SHAP analysis revealed distinct feature contribution profiles: Age ranked as the top contributing feature in the US cohort, whereas hypertension was the primary contributing features in the Chinese cohort.
Conclusion
CVD‐associated feature profiles differ substantially between US and Chinese middle‐aged and elderly CKD patients, supporting the necessity of population‐calibrated classification models rather than applying a unified model across countries.