Introduction to Concepts in Artificial Intelligence and Machine Learning for Pharmacoepidemiologists: Large Language Models
James M. Gwinnutt, Rodrigo de Oliveira, Miriam J. Haviland, Julien H Shippee, Elizabeth Eldridge, Lenon Mendes Pereira, Emily Bratton, Jay Nanavati, Christina DeFilippo MackABSTRACT
Large language models (LLMs) represent a type of generative artificial intelligence (GenAI) that generate and interpret text, with some LLMs able to process multimodal content (e.g., images, audio, video), and can be deployed as part of agents to perform users' tasks. LLMs can perform natural language processing functions such as summarization, translation, and extraction giving them the potential to enhance and scale pharmacoepidemiological and real‐world research by performing tasks such as literature review, data extraction, and medical writing. Despite the growing integration of GenAI tools into research workflows, their technical foundations and methodological implications remain unfamiliar to many pharmacoepidemiologists, who are often responsible for the reliability and accuracy of research that relies on these tools. This paper aims to inform pharmacoepidemiologists about the capabilities and limitations of LLMs to support responsible integration into the field of pharmacoepidemiology, providing an intuitive overview of how LLMs work, focusing on training and text generation, and reviews current and emerging applications in drug effectiveness and safety research and epidemiology. The article addresses challenges associated with LLM use in real‐world evidence generation, including concerns regarding reproducibility, bias, hallucinations, plagiarism, data privacy, and the need for validation.