DOI: 10.1002/jppr.70090 ISSN: 1445-937X

A real‐world analysis of AI chatbot performance for medicines information enquiries

Duncan Yorkston, Tracey Borrie, Paul Chin

Abstract

Background

The provision of medicines information (MI) services requires interpretation and clinical judgement of complex scenarios by pharmacists. To date, few studies have assessed the performance of artificial intelligence (AI) chatbots to assist pharmacists providing MI advice.

Aim

To evaluate the performance and risk associated with two AI chatbots (Microsoft Copilot and Google Gemini) to answer medicines‐related questions.

Method

A sample of 20 questions answered by the local MI service in November 2023 was entered in the two chatbot applications in January 2024 (round 1) and May 2024 (round 2). All questions were preceded with the prompt ‘I'm a pharmacist’. Chatbot responses were evaluated by comparing with a reference answer given by the MI service using a consensus process in the domains of content, patient management, risk of patient harm, and follow up review. Ethical approval was granted by the Canterbury District Health Board Research Office (Reference no: 20311) and the study conforms with the Declaration of Helsinki.

Results

For the 20 questions answered by both chatbots, few of the round 1 responses ( n = 4 for Copilot and n = 2 for Gemini) were considered complete and with adequate information to commence patient management with no risk of harm. Most were incomplete ( n = 13 for Copilot and n = 15 for Gemini) regarding content, but none were high risk of causing harm. In round 1, four responses from Copilot and eight from Gemini were flagged for follow up review. There was no significant difference in performance between chatbots in round 1 (p = 0.68) or between rounds 1 and 2 (Copilot p = 0.25 and Gemini p > 0.99).

Conclusion

Our study results demonstrated the chatbots' responses were typically suboptimal; albeit, a significant minority prompted a follow up to review the chatbot response.

More from our Archive