DOI: 10.1177/23821205261474043 ISSN: 2382-1205

Predicting Step 2 CK Performance Using Automated Feature Selection and Nested Cross-Validation

Padraig Mark Healy, Syed Latifi

Background

Given the recent transition of the USMLE Step 1 exam to a pass/fail system, evaluative emphasis is expected to shift toward the USMLE Step 2 Clinical Knowledge (USMLE-CK) exam, prompting the need for advanced predictive models. In this study, we proposed a multiple linear regression approach incorporating automated feature selection to predict students’ performance on the USMLE-CK exam.

Methods

Our methodology integrated feature selection and model validation processes within a nested cross-validation (CV) framework to predict USMLE-CK scores. We conducted our analysis on data from four undergraduate medical student cohorts (Classes 2020 to 2023 inclusive, n = 117). A range of performance data was included for feature selection, including internal assessment data and National Board Medical Examination (NBME) Clinical Science Subject Exams scores. Utilizing nested CV, we constructed and assessed multiple regression models using four model evaluation metrics: mean CV error, adjusted-R 2 , Mallows’ Cp, and Bayesian Information Criteria.

Results

This led to the selection of a four-predictor model (adjusted-R 2 = 0.68). This model incorporated a combination of NBME exams (Medicine, Neurology and Surgery) and performance in a pre-clinical unit (Gastrointestinal).

Conclusion

Our approach effectively streamlined the process of building a predictive model by merging feature selection with model validation. By creating an interactive, user-friendly dashboard, we empower medical educators to predict students’ USMLE-CK performance. This modeling and deployment approach holds promise for predicting student performances in other assessments.

More from our Archive