DOI: 10.3390/healthcare14152389 ISSN: 2227-9032

Evaluation of Large Language Models in Generating Physical Exercise Rehabilitation Programs for Musculoskeletal Disorders Across Multiple Clinical Scenarios

Yu Fu, Hairui Li, Mingke You, Li Wang, Weizhi Liu, Kai Zhou, Lingcheng Wang, Xi Chen, Gang Chen

Background: Artificial intelligence and large language models (LLMs) are emerging as transformative technologies in medicine. However, their ability to develop physical exercise rehabilitation programs and provide insights into musculoskeletal (MSK) disorders remains underexplored. This study aimed to evaluate the quality and readability of LLM-generated responses to consultation questions addressing various stages of the clinical process encountered by patients with MSK disorders. Methods: This study recruited 50 patients with musculoskeletal disorders and extracted disease-related frequently asked questions from Google search. We developed three clinical scenario-based question types simulating real consultations, which were processed by four LLMs (GPT-3.5-turbo, GPT-4-turbo, GPT-4o, and Claude-3-haiku-20240307). Response quality was assessed by orthopedic specialists and therapists using the DISCERN instrument, and GPT-4o-assisted evaluation was used to extend the assessment after validation with expert ratings. Readability was systematically assessed via six validated indices. Results: Among the 1476 LLM-generated responses, the generated rehabilitation programs demonstrated consistent adherence to the specified query requirements. The DISCERN scores ranged from 26 (poor) to 68 (excellent), with a mean score of 55.60 ± 8.40. The intraclass correlation coefficient (0.68) indicated moderate interrater agreement, and Cronbach’s α showed good internal consistency. The readability scores across six indices indicated that most responses exceeded the recommended reading levels (p < 0.05). Conclusions: LLMs generated moderate-to-high-quality PE rehabilitation recommendations for patients with MSK disorders across simulated consultation scenarios. However, limited supporting materials and suboptimal readability may restrict their effectiveness. With physician oversight, improved readability, and enhanced supplementary resources, LLMs demonstrate considerable potential as supportive tools in orthopedics.

More from our Archive