Evaluating large language model responses to patient questions about acute promyelocytic leukemia: A comparative cross-sectional study
Jing Yin, Xiao Xiong, Sha Ke, Yadan WangObjective
To compare single-generation responses from five publicly accessible large language model chatbot products to patient-facing questions about acute promyelocytic leukemia across safety, accuracy, empathy, information quality, and readability.
Methods
This comparative cross-sectional, product-level evaluation was conducted under public default web-interface conditions. Eighty-two standardized patient- and caregiver-facing questions were each submitted once to ChatGPT, Gemini, Copilot, DeepSeek, and Doubao, yielding 410 single-generation responses. Five experts assessed safety, accuracy, empathy, DISCERN, EQIP, JAMA benchmark score, and GQS. Readability was measured using ARI, CLI, FKGL, GFI, SMOG, and FRES. Inter-rater reliability was quantified with Fleiss’ kappa and ICC(2,1).
Results
Observed safety rates ranged from 85.37% to 93.90%, with no statistically significant between-product differences (P=0.460). ChatGPT showed the highest observed accuracy and information-quality scores and the most favorable readability profile, whereas Doubao showed the lowest observed information-quality scores and the least favorable readability profile. Doubao and DeepSeek showed the highest observed empathy scores. Readability results suggested that most responses were written above recommended patient-facing reading levels. Inter-rater agreement was acceptable to good (Fleiss’ kappa 0.761; ICC 0.756–0.895).
Conclusions
Under the tested public web-interface conditions, most sampled responses met the predefined safety criteria, but unsafe or potentially misleading responses occurred in every product. These findings should be interpreted as single-generation, product-level estimates rather than stable model characteristics and support disease-specific validation, explicit safety review, and clinician oversight before LLM-generated APL information is used for patient education.