Performance Comparison Between Domestic and International Large Language Models in Patient Education for Chinese Patients with Lumbar Disc Herniation: A Cross-Sectional Study
Quan Zhang, Ruizhong Yan, Xiaoliang Liu, Wenzheng Li, Xuhong XueObjective
This study aims to systematically evaluate and compare the performance of different large language models (LLMs) in providing medical advice for lumbar disc herniation (LDH) within the Chinese clinical context, addressing the current lack of standardized assessment tools for orthopedic guidance in China.
Methods
We constructed a standardized dataset covering diagnosis, treatment, surgical risks, and healthcare navigation. Using a unified prompt, we tested eight models: four international (ChatGPT-4o, Gemini-3.0-Pro, Claude-4.0, Grok-4-auto) and four Chinese (DeepSeek, Doubao, Qwen-Max, Kimi). Senior spine surgeons evaluated responses in a blinded manner using the Mika scale (accuracy), modified DISCERN (reliability), EQIP, and GQS. Readability was assessed via the Ludong University Text Grading Platform.
Results
Significant performance differences emerged among models (P < 0.001). DeepSeek and Doubao achieved the highest median accuracy (2.00; IQR 2.00–2.00). We observed a distinct “competency divergence”: international models displayed stronger safety guardrails and logical robustness in complex reasoning, whereas Chinese models demonstrated unique advantages in service contextualization and local resource navigation. Currently, no single model perfectly integrates global medical evidence with local delivery contexts.
Conclusion
DeepSeek, Doubao, and Gemini are the preferred LLMs for LDH patient education in China. While leading domestic models challenge the stereotype of international superiority in medical logic, the general scarcity of “excellent” responses (<25%) indicates LLMs remain auxiliary tools. Future optimization should prioritize Retrieval-Augmented Generation (RAG) anchored in local clinical guidelines to ensure safe deployment.