DOI: 10.1177/20552076261493233 ISSN: 2055-2076

Performance of large language models in managing spinal cord stimulation inquiries: A comparative study

Shengyu Cui, Yuanren Zhai, Haiyue Zhao, Xiaoxu Yu, Qipeng Luo, Jiachen Shan, Jiaxiang Xia, Xin Huang, Zhiqiang Cui, You Wan, Shuiqing Li

Objectives

Spinal cord stimulation (SCS) is a minimally invasive neuromodulation therapy for refractory chronic pain, but its complex treatment paradigm presents significant informational barriers for patients. While large language models (LLMs) show potential as patient-centered educational tools, systematic comparative studies in the neuromodulation field remain insufficient. This study aimed to systematically evaluate and compare the performance of four mainstream LLMs in addressing SCS-related patient inquiries.

Methods

The LLMs were tasked with answering 31 standardized SCS questions across 6 core domains (Basics, Indications, Procedure, Risks, Management, Outcomes). The performance of ChatGPT-5.2, DeepSeek-V3.2, Gemini-3.1 Pro, and Claude-4.6 Sonnet was evaluated across three dimensions: formal quality (EQIP tool), readability (Flesch-Kincaid indices), and content quality (accuracy, comprehensiveness, relevance). All tests were conducted between March and April 2026. Self-correction capacity was further assessed using guided prompts for initially low-rated responses.

Results

DeepSeek achieved the highest mean EQIP score, significantly outperforming ChatGPT (P = 0.0072), Gemini (P = 0.012), and Claude (P = 0.0007). Initial outputs from all four LLMs exhibited low readability, which significantly improved following simplification prompts (all P < 0.001). All models demonstrated high clinical accuracy with no statistically significant differences ( P = 0.5291); Gemini and Claude attained the highest proportions of “Good” ratings (83.87% and 77.42%, respectively). DeepSeek scored highest in comprehensiveness, whereas Claude excelled in relevance. All models showed improvement after self-correction prompts, with ChatGPT correcting all “Poor” responses to “Good”.

Conclusion

Next-generation LLMs demonstrate high clinical accuracy in managing SCS consultation. Despite low readability of original outputs, simple prompting effectively improves accessibility. Each model presents distinct strengths in formal quality, comprehensiveness, relevance, and self-correction. With professional supervision, these LLMs can serve as reliable supplementary educational tools to fill the information gap for SCS patients.