SPERT‐AR: A Unified Framework for Resolving User Story Ambiguities to Improve Story Point Estimation
Waleed Younas, Jing Zhao, Tahreem Iqbal, Muhammad Ali Lodhi, Rui ChenABSTRACT
Objectives
Accurate story point estimation remains a core challenge in Agile software development because it directly affects sprint planning, resource allocation, and delivery predictability. Learning‐based estimators have improved prediction accuracy, but their performance can degrade when user stories contain vague wording, missing conditions, inconsistent phrasing, or overlapping scope. This paper presents SPERT‐AR, a framework designed to evaluate whether ambiguity‐aware refinement of user stories can improve story point estimation while preserving the original repository‐assigned estimation labels.
Methods
SPERT‐AR detects lexical, syntactic, semantic, and pragmatic ambiguity, refines unclear stories using targeted GPT‐4 prompts, and applies expert validation to preserve functional intent. The clarified stories are then used as input to a reinforced transformer estimator. To avoid conflating textual clarification with label revision, the primary evaluation keeps the original repository‐assigned story points unchanged for both original and resolved stories.
Results
Evaluated on ten real‐world GitLab projects, SPERT‐AR reduces mean absolute error by 29.52% and improves standardized accuracy by 23.41% relative to the SPERT baseline under this fixed‐label setting. Beyond the headline accuracy results, the findings show that SPERT‐AR reduces ambiguity while preserving semantic intent, improves estimation stability across baselines, and provides ablation evidence that ambiguity‐aware refinement contributes to the observed improvement beyond model complexity alone.
Conclusion
Overall, the findings suggest that clearer user‐story representations can make story point estimation more reliable, while expert‐adjusted effort interpretations should be treated as secondary evidence rather than primary ground truth.