DOI: 10.1002/ohn.70359 ISSN: 0194-5998

Quality of AI‐Generated Patient Education for Pre‐ and Post‐Operative Tracheostomy Care

Keer Zhang, Lauran K. Evans, Desiree Delavary, Christian Wooten, Minjae Kim, Travis L. Shiba, Natalie Kadin, Joshua D. Feintuch, Dinesh K. Chhetri

Abstract

Objective

To evaluate the accuracy, completeness, clarity, source transparency, and readability of leading AI chatbot responses to patient questions about tracheostomy and to determine whether AI tools can reliably support patient education where high‐quality guidance is critical for safety.

Study Design

Cross‐sectional content analysis.

Setting

Virtual study environment using publicly accessible AI platforms, with expert evaluation conducted via Qualtrics‐based distribution.

Methods

Twelve frequently asked questions about tracheostomy care were identified using search‐listening tools and clinician input, then submitted to 5 AI chatbots — ChatGPT4, Google Gemini 2.0, Microsoft Copilot, DeepSeek V3, and Grok 3 — and to a senior laryngologist. Three blinded laryngologists independently evaluated each response using the Quality Analysis of Medical Artificial Intelligence instrument. Readability was assessed using nine metrics.

Results

Gemini 2.0 achieved significantly higher completeness scores than physician responses ( P  < .001), with DeepSeek and Grok 3 ( P  < .05) also outperforming ( P  < .05). Accuracy did not differ significantly between AI‐ and expert‐generated responses. On average, the AI models outperformed physician in clarity, completeness, and usefulness based on QAMAI scoring ( P  < .05). All AI and expert responses exceeded the NIH‐recommended 6th‐grade reading level, ranging from 10th−13th grade ( P  < .001). Inter‐rater reliability was 78%.

Conclusion

AI chatbots can generate accurate and comprehensive responses to common tracheostomy care questions, demonstrating potential to support patient education. However, they continue to lack guaranteed, verifiable sourcing, and this study did not assess actual patient comprehension of the AI‐generated responses. Future efforts should focus on adapting AI‐generated education materials to meet health literacy standards and evaluating their direct impact on patient understanding and outcomes.

More from our Archive