DOI: 10.36469/001c.165173 ISSN: 2327-2236

Evaluating AI Performance in Systematic Literature Reviews for HEOR: A Case Study

Jing Wang-Silvanto, Mansee Jajoo, He Guo, Rishi Ohri, Judith Nelissen, Carolina Casañas i Comabella

Background

Systematic literature reviews (SLRs) are central to evidence generation in health economics and outcomes research, but traditional SLR workflows are labor-intensive. Artificial intelligence (AI) is increasingly being explored to support evidence synthesis tasks.

Objectives

To describe methods used to assess AI performance in screening and data extraction in SLRs, and to provide recommendations on how to ensure appropriate use of AI in literature reviews.

Methods

We conducted an AI-assisted SLR that mirrored a traditional SLR performed by humans only and assessed the performance of the AI tools employed. AI was used to screen titles/abstracts, screen full texts, and conduct data extraction. Using the traditional SLR as the benchmark, AI performance for title/abstract screening was evaluated in terms of accuracy, recall, and precision at predefined screening milestones, followed by the accuracy assessment of full-text screening by AI. Data items extracted by AI were verified by humans and categorized as correct, incomplete, missing, incorrect, or requiring human checks. The average accuracy rate calculated on the basis of correct data items extracted was used as an indicator of AI performance in data extraction.

Results

During title/abstract screening, accuracy and recall remained high but precision was low, increasing false positives. Full-text screening identified a minority of studies included in the traditional SLR. Mean extraction accuracy was 72.93% (range, 57.69%-88.46%). No data were hallucinated by the AI model used.

More from our Archive