DOI: 10.1155/ijcp/6817934 ISSN: 1368-5031

Machine Learning–Based Classification of Prevalent Hypertension Among Middle‐Aged and Elderly Adults in Chengdu, China

Shan Liu, ShuJie Meng, Yong Li, Xia Wang, Liu Lu, Juan Lian, Caiyi Ma, Rong Zhang, Ru Gao

Background

Hypertension remains a major public health concern worldwide, particularly among middle‐aged and elderly populations; in China, the prevalence among adults has reached 31.6%, and the burden in rapidly urbanizing regions such as Chengdu continues to rise with changing dietary and lifestyle patterns. Identifying individuals with hypertension in hospital‐based check‐up settings is important for timely management. Machine learning, which can capture nonlinear relationships among routinely collected questionnaire‐based variables, may improve the classification of prevalent hypertension in this population.

Objective

This study aimed to assess the prevalence of hypertension among middle‐aged and elderly residents in Chengdu, China, to identify key factors associated with hypertension, and to compare the discrimination and calibration of multiple machine learning algorithms in classifying prevalent hypertension.

Methods

A total of 14,573 individuals aged ≥ 45 years who underwent routine physical examinations at a tertiary general hospital in Chengdu between January and December 2023 were included. Demographic and lifestyle data were collected using standardized questionnaires. Participants were randomly divided into a training set ( n  = 13,116) and a validation set ( n  = 1457) in a 9:1 ratio using stratified sampling. To avoid feature‐selection leakage, univariate analysis, multivariable logistic regression (LR), and collinearity diagnosis were performed exclusively within the training set, yielding 11 independent predictors; the validation set did not participate in any feature‐ or model‐selection step. Four machine learning models, extreme gradient boosting (XGBoost), random forest (RF), LR, and least absolute shrinkage and selection operator (LASSO), were developed, with the synthetic minority oversampling technique applied only to the training folds and hyperparameters tuned by grid search with 10‐fold stratified cross‐validation. A nomogram was constructed based on the LR model. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC), accuracy, sensitivity, specificity, Youden index, and Brier score, with the classification cutoff determined by maximizing the Youden index in the training set.

Results

Among the 14,573 participants, 2823 (19.4%) were diagnosed with hypertension. Eleven independent predictors were identified, including age, sex, employment status, family history of hypertension, satiation level, dietary preferences (high‐salt, spicy, and high‐fat diets), combined diabetes, other cardiovascular disease, and subjective mental stress. In the validation set, all four models showed good discrimination, with AUCs of 0.936 (XGBoost), 0.935 (RF), 0.920 (LR), and 0.920 (LASSO) and accuracy above 0.85; LASSO achieved the highest sensitivity (0.904). Calibration was good for all models (bootstrap Brier scores < 0.15), and decision curve analysis indicated positive net clinical benefit across threshold probabilities of 5%–95%.

Conclusion

Hypertension was common among middle‐aged and elderly adults attending health examinations in Chengdu, China, affecting 19.4% of the participants. All four machine learning models classified prevalent hypertension accurately. XGBoost and RF achieved the highest discrimination, while LR and LASSO were better calibrated. These models may help identify hypertensive individuals in hospital‐based check‐up settings, although their performance still needs to be confirmed in external populations.