DOI: 10.1177/20552076261492798 ISSN: 2055-2076

Multilingual reliability of AI-generated health information on frozen shoulder: A comparative study of three Chinese large language models

Xiaoqing Wang, Min Liu, Weiyi Xu, Chunyuan Cai

Background

Large language models (LLMs) are increasingly used to retrieve medical information and generate patient-readable explanations. However, their performance on musculoskeletal topics across languages remains unclear. Using frozen shoulder as a test case, this study evaluated the bilingual factual accuracy of three Chinese LLMs: DeepSeek, DouBao, and Kimi.

Methods

A 16-item question set based on the 2025 international expert consensus on frozen shoulder was submitted in Chinese and English to three models. Responses were de-identified, randomized, and evaluated by two fellowship-trained orthopedic surgeons. Inter-rater agreement was assessed using Cohen’s kappa (unweighted and quadratically weighted). Inter-model and cross-lingual performance were analyzed using the Kruskal--Wallis test and a repeated-measures question-clustered model with Holm correction.

Results

Inter-rater agreement ranged from moderate to substantial (unweighted κ = 0.449; weighted κ = 0.604), with higher agreement for Chinese than English responses. Chinese-language responses yielded higher mean factual accuracy scores across all models. No significant inter-model differences were observed in either language group (p > 0.05). DeepSeek performed significantly better in Chinese than in English (p < 0.001 and p = 0.004), whereas DouBao and Kimi showed no significant cross-lingual differences.

Conclusions

For all LLMs, the responses had mean factual accuracy scores approaching “accurate,” with infrequent inaccurate responses (6.25%). However, cross-language consistency varied. Prompt language had a model-dependent effect: DeepSeek was significantly more accurate in Chinese, whereas DouBao and Kimi remained consistent. Given the exploratory nature of this study, these findings should be considered as hypothesis-generating rather than as evidence of clinical impact.