A multidimensional evaluation of AI-generated cervical cancer prevention information: Comparing ChatGPT and deepseek
Jinghong Liang, Dong Liu, Wenjuan Liao, Mengyao Huang, Qingjian Ye, Xiaomao LiObjective
This study aimed to evaluate and compare the quality, reliability, readability, and actionability of cervical cancer prevention information generated by ChatGPT and DeepSeek-V3.2 using validated assessment tools.
Methods
A 44-item Frequently Asked Questions (FAQ) bank was developed from international guidelines, Google Trends analysis and expert review. ChatGPT (GPT-5) and DeepSeek-V3.2 generated responses under standardized conditions. Two blinded reviewers evaluated outputs using the Global Quality Score (GQS), DISCERN instrument and the Patient Education Materials Assessment Tool (PEMAT-P). Readability indices and linguistic features (LIWC) were also analyzed. Paired statistical analyses were performed.
Results
DeepSeek generated longer and more complex responses than ChatGPT (p < 0.01). ChatGPT achieved higher scores in evaluated information quality (median GQS: 4.0 [4.0–4.0] vs. 3.0 [3.0–4.0], p < 0.01) and perceived reliability scores (26.5 [24.8–30.0] vs 25.0 [24.0–26.0], p < 0.01), while overall DISCERN ratings were similar and within the “fair” range. Both models demonstrated suboptimal readability and low actionability. Linguistic analysis showed that DeepSeek used more analytical language, whereas ChatGPT exhibited greater social tone and authenticity.
Conclusions
AI chatbots show promise as supplementary tools for cervical cancer prevention education, but limitations in reliability, readability, and actionable guidance persist. Improving evidence transparency, simplifying language, and incorporating structured behavioral guidance are essential to enhance their public health utility.