Automated Software Requirements Elicitation: A Systematic Mapping Study
Safaa Eltahier, Sumaia Mohammed Al-Ghuribi, Mawal A. Mohammed, Imtithal SaeedArtificial intelligence (AI) is transforming requirements elicitation: machine learning, natural language processing (NLP), and large language models (LLMs) now identify software requirements automatically from the textual data that surrounds every project—user feedback, specifications, regulations, and stakeholder transcripts. This paper presents a systematic mapping study of 74 peer-reviewed primary studies on AI-based automated requirements elicitation published between 2021 and 2025, identified from five databases following PRISMA 2020 and classified by AI technique, textual source, elicitation activity, and application domain. The evidence is divided into two equally sized source families—user feedback and agile artefacts versus formal documentation—each coupled to the AI techniques that suit its signal profile. Fine-tuned transformer encoders set the performance ceiling and, task-for-task, still outperform far larger generative models, while LLMs extend elicitation to long regulatory documents, multilingual feedback, and structured outputs. The central finding concerns automation depth. AI identifies requirements with consistently high accuracy (routinely F1 0.8 and above), but automation thins at every subsequent step: 51% of approaches structure what they identify, 23% consolidate them, and only 8% engineer stakeholder validation into the loop. This leaves the steps that turn candidates into agreed requirements largely manual. Benchmark fragmentation (77% custom datasets), thin industrial validation (14%), and skewed non-functional coverage compound this gap. The resulting map gives researchers an evidence-derived agenda for deepening automation, and practitioners guidance on which techniques the evidence supports for each elicitation task and textual source.