DOI: 10.1002/cai2.70077 ISSN: 2770-9191

Evaluating the Performance of Covidence's Artificial Intelligence‐Based Screening Function in Cancer Systematic Reviews

Ashirbani Saha, Xiaomei Yao, Sharan Saravanan, Aditya Misra, Ashley Low, Samya Ali, Mariam Abdelmalek, Shakil Ahmed, Sussman Jonathan

ABSTRACT

Background

Systematic reviews (SRs), often conducted for evidence synthesis, require extensive multistage screening, with Stage‐I title/abstract screening being especially time consuming. Covidence systematic review software, a widely used SR platform, has offered artificial intelligence (AI)‐assisted Stage‐I screening function since 2022. However, this function's performance has not yet been formally compared across heterogeneous, large‐scale oncology‐focused SRs.

Methods

We selected two completed SRs conducted by the Program in Evidence‐based Care (PEBC), Ontario, Canada. The first SR (Axilla‐BC, N  = 8774) supported a clinical practice guideline in breast cancer and the second SR (PET‐utility, N  = 7267) informed the provincial Positron Emission Tomography (PET) steering committee in Ontario. Using each SR, we conducted a simulation‐based study by drawing subsets of 500, 1000, and 2000 articles and running 30, 30, and 10 simulation trials, respectively, to emulate Covidence's AI‐assisted Stage‐I screening. For each trial, we calculated the workload and time savings at 95% and 100% sensitivity for relevant articles and at 100% sensitivity for finally included articles as primary outcomes. Secondary outcomes comprised missed finally included articles at 95% sensitivity for relevant articles. Wilcoxon rank‐sum tests were used to compare outcomes between these two SRs.

Results

When pooling all subset sizes and the N  = 1000 subsets, workload savings at 95% sensitivity were significantly higher for Axilla‐BC than for PET‐utility (median = 37.4% vs. 26.2%, W  = 3156.5, p  = 0.003 and median = 38.3% vs. 20.5%, W  = 641.0, p  = 0.005, respectively). Corresponding time savings were also significantly higher for Axilla‐BC (median = 6.0 vs. 4.2 h, W  = 3089.5, p  = 0.008 and median = 9.6 vs. 5.1 h, W  = 639.0, p  = 0.005). In Axilla‐BC, 20% of the 70 trials missed a total of 18 finally included articles, but only one PET‐utility trial missed one article.

Conclusions

In two large oncology‐SRs, Covidence's AI‐assisted Stage‐I screening function demonstrated meaningful potential to reduce reviewer workload and screening time. Nonetheless, Covidence's performance varied notably between reviews, suggesting a possible association between the results and the characteristics of the SR, including topic, methodological complexity, relevance rate, and inclusion rate. Careful human oversight remains essential to ensure that no potentially finally included studies are missed in evidence synthesis, particularly in nuanced or methodologically complex reviews.