DOI: 10.1002/cpt.70421 ISSN: 0009-9236

Prompt Engineering of Large Language Models for Medication Dose Calculation

Max Weaver, Natasha J. Petry, Jeremy Cauwels, Roxana A. Lupu, Amanda Massmann

This study aims to design and validate a large language model (LLM) framework for systematic extraction of medication dosing across multiple therapeutic classes using electronic health record (EHR) data. Manual dose annotations were completed for 4295 medications across nine therapeutic classes. Five publicly available LLMs (Mistral‐small, Llama 3–70B, Nova lite, DeepSeek, and Claude 3.5 Sonnet) were tested through iterative prompt engineering. Discrepancies between manual and model‐derived doses were assessed using R 2 and classification accuracy. Among 2146 training and 2149 testing samples, Claude achieved high‐performance metrics ( R 2  = 99.32%, accuracy = 92.36%). DeepSeek had R 2 of 97.17% and accuracy of 90.31%. After prompt optimization, both models showed minor improvements (Claude R 2  = 99.34%, accuracy = 92.78%; DeepSeek R 2  = 98.89%, accuracy = 91.01%). Nova lite, Llama, and Mistral‐small performance metrics improved after prompt re‐engineering (Δ R 2 = 9.23%, 12.65%, and 39.03%, respectively). Models had modest agreement with an intra‐class correlation of 0.73 (CI = 0.72–0.74). In testing data, Claude and DeepSeek demonstrated high‐performance metrics R 2  = 99.74%, accuracy = 93.45% and R 2  = 95.58%, accuracy = 91.73%, respectively. Performance improvements were observed across all models following prompt refinement; however, differences may be attributed to model‐specific responses to the prompting strategy used in this study. LLMs show high accuracy for automated medication dose extraction from EHR data. This study can serve as an initial attempt to demonstrate LLMs' capacity to extract medication dosing information from real‐world data. Utilization of LLMs can augment clinicians in clinical practice through medication reconciliation and enhance research endeavors.

More from our Archive