DOI: 10.3390/siuj7040047 ISSN: 2563-6499

Artificial Intelligence Chatbots as Patient Information Sources in Penile Cancer: A Multi-Platform Evaluation

Kunind Oberoi, Sadia Hassan, Dixon Woon, Kapil Sethi

Background/Objectives: Penile cancer carries a disproportionate burden of stigma and delayed presentation, with affected men increasingly turning to artificial intelligence (AI) chatbots as an anonymous information source. This study evaluated the quality, readability, understandability, actionability and clinical accuracy of AI chatbot responses to standardised penile cancer patient queries across six publicly available platforms. Methods: Fourteen standardised questions were submitted to ChatGPT-4o, Gemini, Perplexity, Microsoft Copilot, Claude, and DeepSeek, generating 84 responses. The responses were evaluated using DISCERN (information quality, scored 16–80), Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P) (understandability and actionability, scored 0–100%), and Flesch–Kincaid grade level (reading complexity, recommended threshold ≤Grade 8). Between-platform and between-domain comparisons used the Kruskal–Wallis test with Dunn’s post hoc analysis. Clinical accuracy was assessed against the 2026 European Association of Urology-American Society of Clinical Oncology (EAU-ASCO) Collaborative Guidelines on Penile Cancer using a three-point ordinal scale across 10 guideline-scorable questions (maximum 20 points per platform); the four questions not addressed by the guidelines were scored separately for factual accuracy against authoritative external evidence. Results: No significant between-platform differences were identified for DISCERN (p = 0.796) or PEMAT-P understandability (p = 0.147). All platforms exceeded the 70% understandability adequacy threshold. Perplexity and Copilot demonstrated significantly higher actionability than all other platforms (100% vs. 75%, p < 0.001). No platforms achieved the recommended Grade 8 reading threshold, with Claude generating significantly more complex responses than ChatGPT-4o and Perplexity (grade 11.4 vs. 8.65 and 8.60, p < 0.05). Significant variation was identified across clinical domains for DISCERN (p < 0.001), Flesch–Kincaid grade level (p < 0.001), and word count (p < 0.001), with treatment-related questions achieving the highest information quality but also generating the most complex responses. Clinical accuracy scores ranged from 13/20 (ChatGPT-4o) to 16/20 (Copilot). Two critical errors were identified: Claude and DeepSeek both recommended bleomycin-containing chemotherapy regimens, directly contradicting the EAU-ASCO Strong recommendation against bleomycin due to pulmonary toxicity risk. Responses to the four survivorship and quality-of-life questions were factually accurate against external evidence in 23 of 24 cases. Conclusions: AI chatbot responses to penile cancer patient queries are broadly understandable but consistently fail to meet recommended readability thresholds and provide limited actionable guidance. Two platforms recommended bleomycin-containing regimens against an EAU-ASCO Strong recommendation; no platforms achieved full guideline concordance. Urologists should counsel patients on the limitations of AI chatbots as a health information source.

More from our Archive