DOI: 10.3390/electronics15194456 ISSN: 2079-9292

Mitigating Hallucinations in Finance-Based Multi-Agent Large Language Model Systems

Rashmi Nagpal, Unyimeabasi Usua, Kailey Simons, Sabrina Queipo, Amar Gupta, Vishal Gossain

Large language models are increasingly used for financial question answering, while they are prone to generating hallucinated content. In this research, we propose a multi-signal framework for hallucination detection and mitigation. Our framework combines six signals (entailment, semantic similarity, claim verification, numeric consistency, token overlap and confidence calibration techniques) to identify hallucination type and severity. We evaluated our framework using multiple language models across three financial QA benchmark datasets (FinQA, FinanceQA, and FinDER). Our proposed framework corrects up to 78.4% of the detected hallucinations (fix rate on the hallucinated subset of FinQA; 47.7% on FinanceQA and 45.7% on FinDER), outperforming a type-agnostic mitigation baseline on all three datasets. This highlights that targeted, type-aware hallucination mitigation can significantly improve answer reliability while remaining computationally effective.