DOI: 10.1177/24741264261469005 ISSN: 2474-1264

Ability of Large Language Models to Answer Patients’ Questions and Generate Educational Materials for Uncommon Retinal Conditions

Samuel A. Cohen, Prashant D. Tailor, Adrian Au, Jennifer Vu, Alejandro I. Marin, Pradeep S. Prasad, Julie Kwon, Hamid Hosseini, Jayanth Sridhar

Purpose:

To assess the ability of large language models to accurately and comprehensively respond to frequently asked questions and generate patient education materials related to uncommon retinal conditions.

Methods:

A total of 50 frequently asked questions related to 10 uncommon retinal conditions were input into 3 large language models: ChatGPT-4o1, Google Gemini 2.0 Flash, and Microsoft Copilot (updated January 7, 2025). The accuracy and completeness of responses to frequently asked questions were evaluated by retina specialists using a Likert scale ranging from 1 (very inaccurate/not at all complete) to 5 (very accurate/completely complete), while readability was assessed using validated indices. The large language models were also instructed to generate patient education materials at specific reading levels, which were again evaluated for accuracy, completeness, and readability.

Results:

Responses to patient frequently asked questions were written at mean grade levels of 15.3 ± 1.4 (ChatGPT), 15.7 ± 1.9 (Gemini), and 14.6 ± 1.3 (Copilot), respectively ( P = .02). Mean accuracy scores were 4.02 ± 0.7, 4.16 ± 0.7, and 3.81 ± 0.8. Accuracy scores for large language models-generated patient education materials were 4.70 ± 0.7 (ChatGPT), 4.85 ± 0.4 (Gemini), and 4.30 ± 0.8 (Copilot). When instructed to revise educational materials to improve understandability, all large language models significantly reduced reading levels by 4.8 (ChatGPT), 6.9 (Gemini), and 3.8 (Copilot) grade levels ( P < .001) without compromising accuracy.

Conclusions:

Large language models can accurately respond to patients’ frequently asked questions related to uncommon retinal conditions. Furthermore, large language models can effectively improve the readability of existing education materials for patients with varying levels of health literacy. A deeper understanding of large language model applications may facilitate their integration into clinical practice.

More from our Archive