DOI: 10.1002/wjo2.70153 ISSN: 2095-8811

Performance of Large Language Models in Answering Otolaryngology Clinical Questions in Non‐English Languages

Eugene Oh, Arthur W. Wu, Laura Garcia‐Rodriguez, Zhen‐Xiao Huang, Wei‐Chung Hsu, Jeongkyou Kim, Young Chul Kim, Yi‐Tsen Lin, Hae Chan Park, Hector Andres Perez, Juan San Juan, Gene Liu, Matthew Lee, Dennis M. Tang

ABSTRACT

Background

Large language models (LLMs) such as ChatGPT are being explored for various medical applications, but their performance across languages is not established. Despite the promising results in English from previous studies, it is unclear how ChatGPT performs in other languages. We assessed the ability of ChatGPT‐4 to answer common otolaryngology‐related frequently asked questions in Chinese, Korean, and Spanish in this study.

Methods

Eighteen clinical questions were selected from the American Academy of Otolaryngology‐Head and Neck Surgery Foundation guidelines and ENT Health across six subspecialties. Bilingual clinicians translated the questions to Chinese, Korean, or Spanish and input them into ChatGPT‐4 (OpenAI, San Francisco, USA). The responses were rated on a five‐point Likert scale for accuracy, completeness, and similarity to English reference answers. A linear mixed‐effects model with language as a fixed effect and random intercepts for question and rater was used to compare performance across languages. Inter‐rater reliability was assessed using intraclass correlation coefficients and Krippendorff's α . Subjective evaluations were also collected.

Results

ChatGPT‐4's performance differed significantly by language in all assessment domains. Responses in Spanish scored highest throughout and most frequently were characterized as clear and organized. Korean answers were generally accurate but tended to be more conservative and less detailed. Chinese answers showed greater variability, with reviewers noting a mix of clear and less accurate outputs.

Conclusion

The clinical performance of ChatGPT‐4 differed across Chinese, Korean, and Spanish. These findings underscore the importance of ongoing evaluation and refinement of LLMs to ensure they provide reliable multilingual support in clinical practice.

More from our Archive