Evaluation of responses generated by chatGPT and deepseek to drug information enquiries in the hospital setting
Joyce H.S. You, Wai-li Lim, Alan C.L. Wong, Meyone Y.T. Ng, Johnny S.H. Yuen, Bilvick B.W. Tai, Kathy Chow, Teddy T.N. LamBackground
This study aims to compare the performance of ChatGPT (version GPT-4.1) and DeepSeek (R1) in key categories of drug information (DI) enquiries in the Hong Kong hospital setting.
Methods
78 questions (on 13 DI categories, with 6 questions per category) were retrieved from DI enquiry logs in two teaching hospitals. ChatGPT and DeepSeek were separately used to answer each DI question and generated total 156 responses. Five pharmacists individually evaluated each response on 8 domains using a 5-point Likert scale to generate 6240 assessment score.
Results
The grand average value of all scores was 3.79±1.26 (out of 5 as maximum). By DI category, the highest total average score was “drug use in pregnancy and lactation” (4.15±1.05), and the lowest score was “pharmacoeconomics” (3.06±1.33). By response domain, the highest total average score was “language/readability” (4.62±0.64), and lowest score was “absence of unclear/misleading information” (3.32±1.38). The average overall score of DeepSeek (3.84 ± 1.20) was higher than ChatGPT (3.74 ± 1.31) by 0.10 (p<0.001).
Conclusion
ChatGPT and DeepSeek-generated responses to queries in 13 DI categories showed the highest overall rating in “drug use in pregnancy and lactation”, and the lowest overall rating in “pharmacoeconomics”. The overall rating of response domains was highest in “language/readability” and lowest in “absence of unclear/misleading information”. The overall performance of DeepSeek was higher than ChatGPT by a modest difference.