A Multidimensional Evaluation of Large Language Model Responses to the 2024 ESC Hypertension Guidelines: A Comparative Study
Ferit Böyük, Aysun Karahan Gün, İsmail Polat Canbolat, Emre Özmen, Harun Akarsu, Halime Tanrıverdi, Bilal Cuğlan, Kemal Türker UlutaşLarge language models (LLMs), including ChatGPT, Gemini, and DeepSeek, are increasingly used in medicine; however, their performance across multiple clinically relevant domains remains incompletely understood. This cross-sectional comparative study evaluated 225 responses generated by three LLMs to 75 clinical questions selected from the 2024 European Society of Cardiology (ESC) Hypertension Guidelines. Responses were independently assessed by two board-certified cardiologists across five predefined domains: accuracy, clinical relevance, completeness, absence of bias and misinformation, and consistency. No statistically significant differences were observed among the three models in accuracy, clinical relevance, completeness, or absence of bias and misinformation (all p > 0.05). A significant difference was identified in response consistency (p = 0.032), with post hoc analysis demonstrating a significant difference between ChatGPT and Gemini. Overall appropriateness scores did not differ significantly among the three LLMs (p = 0.227). These findings suggest that, although overall performance was comparable, response consistency represents an additional dimension that should be considered when evaluating LLMs for guideline-based clinical applications. Future studies incorporating broader clinical scenarios and updated LLM versions are warranted to further define their role in clinical decision support.