An Agentic AI-Aided Review of Large Language Model Applications in Literature Reviews: A Seven-Layer Architecture
Emmanuel A. Merchán-Cruz, Ioseb Gabelaia, Shwe Soe, Irina Yatskiv, Dmitry PavlyukThe growing volume of scientific output and the pace of AI advancement create a dual challenge: traditional systematic reviews take twelve to eighteen months to complete, risking obsolescence before publication, while the technology needed to accelerate them is itself advancing faster than it can be reviewed. This work addresses both sides of that challenge simultaneously—presenting an agentic systematic literature review of AI-assisted reviewing, conducted using the same agentic pipeline it evaluates, across 122 primary studies published between 2022 and 2026. LLM performance proves to be sharply stage-specific: screening and structured extraction are tractable under human supervision, quality appraisal remains near-chance, and no published pipeline meets the bar for unsupervised end-to-end deployment. The review proposes a seven-layer reference architecture and identifies three prerequisite investments for production readiness: a shared benchmark, validated risk-of-bias performance, and a mandatory AI reporting standard.