Modeling Reliability and Validity Across OSCE Station Numbers in Final MBBS Exam
Alok Kumar, Kandamaran Krishnamurthy, Shastri Motilal, Michael H Campbell, Joanne Paul-Charles, Maritza Fernandes, Kenneth Connell, Euclid Morris, Maisha Emmanuel, Bidyadhar Sa, Md Anwarul Azim MajumderBackground: The Objective Structured Clinical Examination (OSCE) has been the backbone of assessment in medical education, and the number of stations required to achieve sufficient reliability in high-stakes examinations is an important consideration for resource-limited programs. Methods: The correlation between the number of stations and psychometric performance was simulated using post hoc resampling of the 17-station Final MBBS Medicine and Therapeutics OSCE across three campuses of the University of the West Indies. Blueprint coverage was maintained in balanced subsets of six, eight, ten, and twelve stations. For subsets and the full exam, we compared failure rates, internal consistency, generalizability, and correlation of scores. Results: Failure rates increased as the number of stations decreased; however, none of the differences between subset and full-exam performance were significant. Station count was positively related to internal consistency using alpha (α ≈ 0.32–0.61 at six stations; α ≈ 0.64–0.70 at 17 stations) and G-coefficients (≈0.18–0.42 at six and ≈0.36–0.70 at 17 stations). However, all reliability coefficients were <0.80. Strong internal consistency was maintained for all subsets, even at six stations (r ≥ 0.79), and approached 1.0 for the 17-station subset. There was a subset size effect on mean values for Campuses 1 and 2 but not for Campus 3. Conclusions: Shorter subset exams maintained strong internal consistency. Reliability at ten–twelve stations approached that of the 17-station total. Intercampus variability highlighted the need for examiner training, station quality checks, and improved standard setting.