Comparative Analysis of Regularised Logistic Regression and Random Forest Models for In-hospital or 30-day Post-Discharge Mortality Prediction within the Hospital Standardised Mortality Ratio Framework
Mohd Kamarulariffin Kamarudin, Sarbhan Singh Lakha Singh, Nuur Hafizah Md Iderus, Sumarni Mohd Ghazali, Lonny Chen Rong Qi Ahmad, Nur'ain Mohd Ghazali, Mohamad Nadzmi Md Nadzri, Asrul Anuar Zulkifli, Mei Cheng Lim, Chien Huey Teh, Mohd Azahadi OmarObjectives: The hospital standardised mortality ratio (HSMR) is the ratio of the observed number of hospital deaths to the expected number of deaths, with the latter estimated using statistical models that adjust for available case-mix factors. This study aimed to develop and validate in-hospital or 30-day post-discharge mortality prediction models for 40 diagnosis groups within the HSMR framework, using penalized logistic regression (pLR) and random forest (RF), and to compare the performance of these two approaches.Methods: We analysed 1,144,890 hospital admissions from 14 Malaysian state hospitals between 2012 and 2016. Separate models were developed for each diagnosis group using nine administrative features, including age, comorbidities, and admission category. Model performance was evaluated using mean Brier scores and Cstatistics across multiple bootstrapped datasets to obtain less biased performance estimates. Aggregate expected mortality counts were also compared with observed counts.Results: The overall observed mortality rate was 10.2%. The pLR models consistently showed better discrimination and calibration than the RF models, with lower Brier scores and higher C-statistics across the 40 diagnosis groups. On average, the C-statistic for pLR exceeded that for RF by 0.062. Although the RF model produced aggregate mortality predictions that were numerically closer to the observed counts, it showed high variance and poorer probabilistic calibration than pLR.Conclusions: The pLR model tended to underestimate mortality more than RF but still demonstrated better calibration and discrimination, making it the preferable model for HSMR analysis in this dataset.