DOI: 10.60118/001c.162882 ISSN: 2691-6541

Evidence on Demand: OpenEvidence versus Established LLMs on the Orthopaedic In-Training Examination

Asim A. Khan, Shaan S. Lalvani, Sam Pourarbab, Lord J. Hyeamang, Ryan M. Lew, Zaamin B. Hussain, Gregory P. Nicholson, Grant E. Garrigues

Background

Large language models (LLMs) are increasingly used in medical education, yet head-to-head comparisons of contemporary multimodal and clinically oriented (retrieval-grounded) models on orthopedic in-service examinations remain limited.

Methods

A cross-sectional comparative study was conducted using 265 multiple-choice questions from the publicly available 2014 Orthopaedic In-Training Examination (OITE). Questions were categorized by AAOS subspecialty and by cognitive taxonomy (T1 recall, T2 interpretation, T3 reasoning) and stratified into text-only (n=111) versus image-based items (n=164). ChatGPT-5 Plus and Gemini 2.5 Pro were evaluated on the full dataset, while OpenEvidence was evaluated on the text-only subset. Accuracy was scored against the official answer key and comparisons were done using McNemar and Cochran Q tests.

Results

Overall accuracy on the full 2014 OITE was 80.0% for Gemini 2.5 Pro and 78.1% for ChatGPT-5, with no statistically significant difference (p=0.59). Across the text-only subset, Gemini 2.5 Pro and OpenEvidence each scored 84.7% versus 81.1% for ChatGPT-5 (p=0.71), and OpenEvidence demonstrated strong T3 reasoning performance (82.7%). Subspecialty performance varied without statistically significant differences. LLMs demonstrated highest accuracy in basic science and the largest numerical discrepancy was in Oncology (Gemini 77.3% vs ChatGPT-5 54.5%).

Conclusion

A retrieval-grounded, evidence based LLM (OpenEvidence) performed comparably to multimodal LLMs on text-based orthopaedic in-training examination questions at or above senior resident PGY-5 benchmark-level performance with no significant differences in overall accuracy. These findings suggest that OpenEvidence may provide educational utility comparable to general-purpose models while offering source transparency.