Evaluating AI Performance in Systematic Literature Reviews for HEOR: A Case Study
Jing Wang-Silvanto, Mansee Jajoo, He Guo, Rishi Ohri, Judith Nelissen, Carolina Casañas i ComabellaBackground
Systematic literature reviews (SLRs) are central to evidence generation in health economics and outcomes research, but traditional SLR workflows are labor-intensive. Artificial intelligence (AI) is increasingly being explored to support evidence synthesis tasks.
Objectives
To describe methods used to assess AI performance in screening and data extraction in SLRs, and to provide recommendations on how to ensure appropriate use of AI in literature reviews.
Methods
We conducted an AI-assisted SLR that mirrored a traditional SLR performed by humans only and assessed the performance of the AI tools employed. AI was used to screen titles/abstracts, screen full texts, and conduct data extraction. Using the traditional SLR as the benchmark, AI performance for title/abstract screening was evaluated in terms of accuracy, recall, and precision at predefined screening milestones, followed by the accuracy assessment of full-text screening by AI. Data items extracted by AI were verified by humans and categorized as correct, incomplete, missing, incorrect, or requiring human checks. The average accuracy rate calculated on the basis of correct data items extracted was used as an indicator of AI performance in data extraction.
Results
During title/abstract screening, accuracy and recall remained high but precision was low, increasing false positives. Full-text screening identified a minority of studies included in the traditional SLR. Mean extraction accuracy was 72.93% (range, 57.69%-88.46%). No data were hallucinated by the AI model used.