DOI: 10.1002/tre.70013 ISSN: 2044-3730

Evaluating the Role of AI Chatbots in Patient Education for Benign Scrotal Surgeries

Darcy Noll, Thomas Milton, Peter Stapleton, Kathryn Sharley, Richard Hoffmann

ABSTRACT

Patient comprehension of surgical information is often limited by medical jargon, health literacy barriers, and time constraints. Large language models (LLMs) such as ChatGPT‐5, Claude‐4, and Google AI search offer interactive context specific dialogue that may aid to overcome these limitations. To date, no study has assessed the ability of LLMs to provide patient education for hydrocelectomy, spermatocelectomy and vasovasostomy. The objective of this study was to assess the quality, readability, and understandability of LLM‐generated content for these procedures. The most common patient queries for each procedure were identified using Google Trends and educational materials from professional healthcare organisations. Questions were standardised into lay language and posed to ChatGPT‐5, Claude‐4 and Google AI search. Outputs were assessed using validated tools: DISCERN‐AI for quality, Global Quality Score (GQS) for quality and information flow, and PEMAT‐AI for understandability. Readability was assessed with Flesch Reading Ease and Flesch‐Kincaid Grade Level (FKGL). All models produced high‐quality responses (median DISCERN‐AI 13–14) and satisfactory understandability (median PEMAT 71%–86%). Median GQS was 4, indicating good information flow, though no responses achieved the maximum rating. Readability was consistently above population literacy levels for all three models (ChatGPT FKGL = 11.2, Claude FKGL = 14.2, Google FKGL = 12.3). When asked to assess the readability of their own outputs, the readability was overestimated by the LLMs in every instance. LLMs provide high quality, understandable and mostly accurate patient educational content on elective scrotal procedures for benign pathology. However, outputs are written at excessively high reading levels, potentially limiting accessibility. Future improvements in readability and fact‐checking are essential before widespread patient adoption.

More from our Archive