DOI: 10.1111/exsy.70437 ISSN: 0266-4720

Large Language Models and Agentic AI for Vulnerability Management: A Systematic Review

Yasamin Akrami, Malek Malkawi, Reda Alhajj

ABSTRACT

Large Language Models (LLMs), pretrained language and code models and agentic artificial intelligence are increasingly being investigated across the vulnerability‐management lifecycle. However, the evidence remains fragmented across code‐level prediction, vulnerability‐intelligence enrichment, exploit‐oriented evaluation, automated repair and bounded defensive action. Following PRISMA 2020 guidance, this systematic review searched six scholarly databases, supplemented the initial search with an update search conducted on 22 August 2026 and backward and forward citation searching, and identified 53 eligible peer‐reviewed primary studies published between 2023 and 2026. The studies were organized into four application areas: vulnerability discovery and detection; vulnerability intelligence and prioritization; threat context, exploitation and operational hunting; and remediation, response and autonomous defence. Model‐centred studies provide the most standardized benchmark evidence, but their reported performance remains sensitive to dataset construction, class imbalance, leakage, code context, language coverage and cross‐dataset transfer. More recent work increasingly combines retrieval, structured program or security knowledge, external tools, multi‐agent coordination and verification; however, only 15 of the 53 studies (28.3%) implemented or directly evaluated systems meeting this review's operational definition of genuinely agentic behaviour through adaptive evidence acquisition, iterative tool use, feedback‐driven control, coordination, or bounded action. Exploit‐oriented and autonomous‐defence evaluations provide useful capability evidence but should not be interpreted as direct evidence of operational vulnerability‐management autonomy. Across remediation and response studies, execution success alone does not establish vulnerability removal, semantic correctness, operational safety, or production readiness. To separate technical results from operational credibility, the review applies an evidence‐maturity framework and analyses recurring forms of data, evaluation, grounding, verification and governance debt. The findings support bounded and auditable use of LLMs within evidence‐grounded vulnerability‐management workflows rather than unrestricted autonomous operation.