A Locally Executed Agentic Artificial Intelligence Framework for Deduplication, Screening, and Structured Data Extraction in Spine Surgery Systematic Reviews
Sathish Muthu, Dhibin Vikash Kolarpatti Ponnusamy, Vibhu Krishnan Viswanathan, Sathish Kumar Rajappan Chandra, Khan SharunAbstract
Systematic reviews and meta-analyses remain time-consuming and labor-intensive. We developed and validated a locally executed agentic artificial intelligence (AI) framework for deduplication, screening, and structured data extraction in systematic reviews, reported following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA)-Transparent Reporting of AI in Comprehensive Evidence Synthesis (trAIce) guidelines. A multiagent pipeline of specialized agents for deduplication, title/abstract screening, structured data extraction, and verification was executed entirely locally to ensure data governance and reproducibility. Human-in-the-loop validation compared outputs against dual independent reviewers across 3 spine surgery systematic reviews, assessing accuracy, inter-rater agreement (Cohen’s κ), time savings, and clinically critical error rates. Across 6,214 records, deduplication achieved near-perfect agreement with human reviewers (κ = 0.98), title and abstract screening yielded higher concordance than human screening (κ = 0.91) while reducing full-text review volume by 83%, and structured data extraction reached substantial agreement (κ = 0.87). The framework reduced reviewer time by 91.1% (90.8%-91.4%), a mean saving of 19.5 hours per review (p < 0.001). Clinically critical discrepancies were rare (<1%) and traceable, with no fabricated or hallucinated data introduced. A locally executed agentic AI framework has the potential to deliver accurate, efficient, and secure automation of systematic review tasks with human oversight, offering a reproducible pathway for trustworthy evidence synthesis under PRISMA-trAIce standards.