DOI: 10.1002/wjo2.70158 ISSN: 2095-8811

Vocal Cord Dysfunction: Evaluating the Utility of AI Large Language Models for Patient Education

Kyle Cook, Phil Tseng, David Ahmadian, Natalie Demirjian, Vicki Liu, Troy Weinstein, Helena Yip

ABSTRACT

Objective

To compare provider preferences for patient education materials generated by OpenEvidence, ChatGPT 5 Extended Thinking, and a laryngologist with fellowship training in response to common patient questions about vocal cord dysfunction (VCD).

Study Design

Cross‐sectional survey study.

Methods

A laryngologist compiled four common patient questions about VCD and authored expert responses. Equivalent responses for patients were generated using OpenEvidence and ChatGPT 5 Extended Thinking with a standardized role prompt, each in a new chat session, and all outputs were limited to four sentences. Responses were de‐identified and presented in a REDCap survey. Forty‐five healthcare providers ranked responses based on perceived medical accuracy, clarity, and ability to address patient concerns. Preferences were analyzed using chi‐square goodness‐of‐fit and Friedman testing with post hoc pairwise comparisons.

Results

Among the 45 respondents, the most represented specialties were otolaryngology ( n  = 17) and allergy/immunology ( n  = 12), and most were attending physicians ( n  = 31). ChatGPT 5 Extended Thinking received the most first choice selections for three of the four questions, whereas OpenEvidence and ChatGPT performed similarly on the definition question. Analyses of full rankings confirmed significant differences among sources for all four questions (all Friedman p  ≤ 0.001): responses generated by AI were preferred over the laryngologist response for the definition, diagnosis, and treatment process questions, while ChatGPT was preferred over both comparators for the etiology question. Otolaryngology respondents were more likely than non‐otolaryngology respondents to rank the laryngologist response first for the definition question only.

Conclusion

Responses generated by AI were frequently preferred over a single expert comparator for common VCD questions, with ChatGPT 5 Extended Thinking performing best overall and OpenEvidence showing similar performance on several items. These findings support further evaluation of LLMs as adjunctive tools in laryngology patient education, while highlighting the need for clinician oversight, assessment of individual model performance, and studies of patient outcomes.

More from our Archive