DOI: 10.4103/wkmj.wkmj_81_26 ISSN: 2707-6180

Machine Learning-based Screening for Elevated Cholesterol Using Noninvasive Population-level Predictors

Akmaral Baspakova, Sanimgul S. Sambayeva, Kulyash R. Zhilisbayeva, Roza Suleimenova, Gulden Yelgondina, Akmeiir E. Kaliyeva, Aigerim A. Umbetova

A
BSTRACT

Introduction:

Machine-learning approaches can capture complex relationships among routinely collected risk factors and support screening-oriented identification of individuals with a higher probability of elevated cholesterol.

Methods:

We developed a supervised machine-learning model using noninvasive predictors based on a population-based cross-sectional dataset from West Kazakhstan (World Health Organization STEPwise approach, 2021–2022). The final sample included 4788 adults aged 18–69 years, with an elevated cholesterol prevalence of 11.8% (565/4788). Elevated cholesterol was defined using standardized laboratory measurements and encoded as a binary outcome for labeling. Predictors included sociodemographic, behavioral, and anthropometric variables (age, sex, smoking, alcohol use, physical inactivity, body mass index (BMI), and waist circumference), processed through a unified preprocessing pipeline. A class-weighted Random Forest model was trained using a stratified train–test split to preserve prevalence. Performance was primarily evaluated using precision–recall (PR) analysis under sensitivity-oriented thresholds, and model interpretability was assessed using SHapley Additive exPlanations (SHAP).

Results:

The model generated probability-based estimates suitable for screening applications. Under class imbalance, PR analysis showed modest discrimination above baseline, with an average precision of ~0.18 versus a baseline of ~0.12. SHAP results indicated age as the strongest predictor, followed by adiposity-related measures (waist circumference and BMI). Behavioral factors (physical inactivity, smoking, and alcohol use) contributed less consistently.

Conclusion:

An explainable machine-learning model using routinely collected variables can support screening-oriented risk stratification for elevated cholesterol. Designed as a decision-support and prescreening tool rather than a diagnostic system, the framework was implemented as an interactive web-based application to enable transparent, real-time risk estimation for preventive and educational use.

More from our Archive