Multilingual reliability of AI-generated health information on frozen shoulder: A comparative study of three Chinese large language models
Xiaoqing Wang, Min Liu, Weiyi Xu, Chunyuan CaiBackground
Large language models (LLMs) are increasingly used to retrieve medical information and generate patient-readable explanations. However, their performance on musculoskeletal topics across languages remains unclear. Using frozen shoulder as a test case, this study evaluated the bilingual factual accuracy of three Chinese LLMs: DeepSeek, DouBao, and Kimi.
Methods
A 16-item question set based on the 2025 international expert consensus on frozen shoulder was submitted in Chinese and English to three models. Responses were de-identified, randomized, and evaluated by two fellowship-trained orthopedic surgeons. Inter-rater agreement was assessed using Cohen’s kappa (unweighted and quadratically weighted). Inter-model and cross-lingual performance were analyzed using the Kruskal--Wallis test and a repeated-measures question-clustered model with Holm correction.
Results
Inter-rater agreement ranged from moderate to substantial (unweighted κ = 0.449; weighted κ = 0.604), with higher agreement for Chinese than English responses. Chinese-language responses yielded higher mean factual accuracy scores across all models. No significant inter-model differences were observed in either language group (p > 0.05). DeepSeek performed significantly better in Chinese than in English (p < 0.001 and p = 0.004), whereas DouBao and Kimi showed no significant cross-lingual differences.
Conclusions
For all LLMs, the responses had mean factual accuracy scores approaching “accurate,” with infrequent inaccurate responses (6.25%). However, cross-language consistency varied. Prompt language had a model-dependent effect: DeepSeek was significantly more accurate in Chinese, whereas DouBao and Kimi remained consistent. Given the exploratory nature of this study, these findings should be considered as hypothesis-generating rather than as evidence of clinical impact.