DOI: 10.1177/20552076261486992 ISSN: 2055-2076

Response quality of large language models on periprosthetic fractures: A comparative evaluation of ChatGPT, Gemini, and Claude

Buğra Can, Anıl Agar, İrfan Arslan

Background

Large language models (LLMs) are increasingly used for patient education, yet their performance in high-risk orthopaedic conditions remains uncertain. This study compared the accuracy, information quality, safety, and readability of ChatGPT, Gemini, and Claude responses to questions about periprosthetic fractures.

Methods

Fifty candidate questions were reduced to 25 standardized patient-oriented questions. On December 27, 2025, each question was submitted once, in a separate session and without follow-up prompts, to ChatGPT GPT-5.2, Gemini 3 Flash, and Claude Sonnet 4.5 using default web-interface settings. Seventy-five anonymized responses were independently assessed by two orthopaedic residents and one orthopaedic trauma specialist using the Mika accuracy score, DISCERN, harm flag, word count, and readability indices.

Results

Mika scores differed significantly among models (p<0.001; Kendall’s W=0.50), with Gemini performing better than ChatGPT and Claude. DISCERN scores also differed significantly (p=0.002; Kendall’s W=0.25), again favoring Gemini. Harm-rating agreement was low (ICC(2,1)=0.16), precluding a definitive comparison of model safety. No response received a harm score of 2. Gemini produced the longest responses, whereas Claude required the highest reading level. Inter-rater agreement was moderate to good for Mika and DISCERN.

Conclusions

LLM performance on periprosthetic fracture questions varied by model. Gemini 3 Flash showed higher accuracy and information quality, whereas Claude Sonnet 4.5 produced less readable responses. Because safety findings were limited by low inter-rater agreement and a single-query design, LLM-generated information should not be used alone without physician oversight, topic-specific safety review, transparent sourcing, and readability optimization.