Ability of Large Language Models to Answer Patients’ Questions and Generate Educational Materials for Uncommon Retinal Conditions
Samuel A. Cohen, Prashant D. Tailor, Adrian Au, Jennifer Vu, Alejandro I. Marin, Pradeep S. Prasad, Julie Kwon, Hamid Hosseini, Jayanth SridharPurpose:
To assess the ability of large language models to accurately and comprehensively respond to frequently asked questions and generate patient education materials related to uncommon retinal conditions.
Methods:
A total of 50 frequently asked questions related to 10 uncommon retinal conditions were input into 3 large language models: ChatGPT-4o1, Google Gemini 2.0 Flash, and Microsoft Copilot (updated January 7, 2025). The accuracy and completeness of responses to frequently asked questions were evaluated by retina specialists using a Likert scale ranging from 1 (very inaccurate/not at all complete) to 5 (very accurate/completely complete), while readability was assessed using validated indices. The large language models were also instructed to generate patient education materials at specific reading levels, which were again evaluated for accuracy, completeness, and readability.
Results:
Responses to patient frequently asked questions were written at mean grade levels of 15.3 ± 1.4 (ChatGPT), 15.7 ± 1.9 (Gemini), and 14.6 ± 1.3 (Copilot), respectively (
Conclusions:
Large language models can accurately respond to patients’ frequently asked questions related to uncommon retinal conditions. Furthermore, large language models can effectively improve the readability of existing education materials for patients with varying levels of health literacy. A deeper understanding of large language model applications may facilitate their integration into clinical practice.