TGAR-HB: A Taxonomy-Guided Retrieval Framework for Java Fault Localization via Evolutionary Testing and LLMs
Fei Xu, Chenhan Wu, Rui XuAccurate fault localization plays a critical role in reducing software maintenance costs. Although LLMs exhibit strong code comprehension capabilities, their effectiveness remains limited when applied to large-scale software projects, primarily due to the lack of runtime execution information and insufficient retrieval precision. This paper proposes a novel fault localization approach that integrates the evolutionary testing tool EvoSuite with LLMs, enhanced by a mechanism termed TGAR-HB (taxonomy-guided adaptive retrieval with hierarchical backtracking). The framework first leverages EvoSuite to automatically generate unit tests, thereby capturing dynamic runtime context. It then introduces a confidence-driven hierarchical backtracking mechanism to adaptively refine the retrieval process. Specifically, the input fault information is categorized into three hierarchical levels—code domain, functional module, and fault type—enabling more precise matching with historical fault-localization experiences stored in a knowledge base. This design aims to reduce the risk of hallucination by grounding LLM reasoning in retrieved knowledge. Experimental results on the Defects4J benchmark demonstrate that the proposed method significantly outperforms traditional spectrum-based fault localization (SBFL) techniques such as Ochiai, as well as plain LLM and standard RAG baselines. Notably, the DeepSeek + TGAR-HB combination achieves the best performance, with a Top-1 accuracy of 34.73% and a Top-10 hit rate of 78.18%, while maintaining a low Mean Average Rank (MAR) of 3.13. Furthermore, ablation studies confirm the effectiveness of the TGAR-HB framework for code retrieval. Overall, this work presents a new approach for automated software debugging by synergistically combining dynamic testing with LLM-based semantic reasoning.