DOI: 10.1177/20552076261490211 ISSN: 2055-2076

Evaluation of large language model-based chatbots in hearing aid consultations: A comparative study of experts and non-experts

Yuseon Byun, Donghyeok Lee, Chanbeom Kwak, Chul Young Yoon, Young Joon Seo

Objective

Hearing loss patients frequently seek reliable information and guidance regarding hearing aids. This study evaluated the clinical suitability and user preference of general-purpose large language models (LLMs) in hearing aid consultations to identify the current limitations of generic models and the necessity for domain-specific solutions.

Methods

A dataset of 1,632 real-world inquiries was refined into 33 representative items across seven audiological categories through a blind review and unanimous consensus by a panel of experienced otolaryngologists and audiologists. Three LLMs (ChatGPT, HyperCLOVA X, and MaumAI) generated responses to these items. To establish a clinical reference standard, the author panel used a unanimous consensus protocol to select the single most accurate response from ten repetitions per model. The accuracy and preference of these responses were evaluated on a 3-point scale by 20 hearing healthcare experts and 33 non-experts (patients/guardians).

Results

Both expert and non-expert groups showed a statistically significant preference for ChatGPT (LLM 1) regarding overall accuracy and preference. However, consistency within and between the two groups was notably low (Kendall’s W < 0.07), revealing a critical ‘perspective gap’. While experts prioritized clinical precision and technical specificity, non-experts favored responses that were practical, empathetic, and easy to understand.

Conclusions

Current general-purpose LLMs struggle to simultaneously satisfy rigorous medical standards and user-centric readability, highlighting the limitations of a “one-size-fits-all” approach. To effectively bridge the health literacy gap, future research should focus on developing a customized Small Large Language Model (sLLM) trained on curated audiological data. Such a domain-specific model is essential to provide both medically reliable and patient-friendly digital hearing healthcare.