Artificial Intelligence Resources for the Screening of Titles and Abstracts in Systematic Reviews: A Scoping Review
Ana M. Barragán, Sara Elena Ortiz Bonett, Eliana‐Isabel Rodríguez‐Grande, Alvaro David Orjuela‐Cañón, Oscar J. Perdomo, Guillermo Sánchez‐VanegasABSTRACT
Introduction
Artificial intelligence (AI) is a branch of technology enabling machines to emulate complex human skills; it can also entail problem‐solving using bioinspired methods. It is used for automating systematic literature reviews (SLR), that is, defining a clinical question, locating relevant literature, preliminary screening, study evaluation, data extraction and analysis. Title and abstract screening is one of the most time‐consuming and error‐prone phases involved in developing a systematic review. While AI promises to expedite this process, adopting it faces challenges due to concerns about compatibility and transparency. This review aims to identify current evidence concerning AI use during preliminary SLR reference screening; it describes characteristics such as the different metrics used for reporting performance and how the different algorithms, pipelines, workflows or web applications are validated. AI resource users' reflections regarding SLR screening automation have also been summarized.
Methods
A scoping review was conducted following Joanna Briggs Institute's (JBI) methodology. Its objective was to identify existing evidence regarding the use of AI resources for title and abstract screening automation. Searches were limited to articles published between 2019 and 2026. The review included primary studies reporting the development, assessment, validation, or real‐world use of AI resources for screening automation, as well as systematic reviews and articles reporting experiences or recommendations for their use. Two types of data were extracted: (1) from primary studies—characteristics of AI resources and, where applicable, recommendations for their use; (2) from systematic reviews and experience‐based articles, recommendations for the use of AI resources. Results included frequency descriptions, tables, figures, and a decision flowchart reflecting the number of references and articles retrieved, excluded, or included in the final analysis.
Results
A total of 174 unique studies published between 2019 and 2026 were included in this scoping review. These were grouped into web applications (43%), model comparisons (32%), generative models (6%), pre‐trained models (3%) or pipelines/workflows (15%) used for title and abstract screening in systematic literature review (SLR). Most studies came from North America. Evaluating these tools often relied on retrospective comparisons with human reviewers' work (63%), sensitivity ( n = 60), and specificity ( n = 62) being the most reported metric for criterion assessment and Work Saved over Sampling ( n = 28) being the most reported metric for assessing their utility. Considerations concerning AI resource use focused on the need for standardized evaluation metrics, stopping criteria, study design and the data sets used, resource characteristics facilitating usability, best practice and future research areas, with the persistence of the human component in the process ( n = 26) being the most pressing recommendation.
Conclusion
The findings indicated substantial heterogeneity regarding the types of AI resources used, considerable variation concerning the metrics used for reporting performance, differences in how such metrics are defined and a clear need for standardizing reporting methods, study designs and related procedures. Although AI technologies will continue to evolve, maintaining a clear and consistent framework for interpreting research on AI resources for automating title and abstract screening can support understanding their level of maturity and facilitate informed decision‐making by users.