DOI: 10.1145/3839536 ISSN: 2475-1421

LLM-Based Alarm Resolution Guided by Bayesian Program Analysis

Yifan Zhang, Yuanfeng Shi, Haoran Lin, Yingfei Xiong, Xin Zhang

Abstract-interpretation-based static analyzers often report large numbers of alarms due to over-approximation. Although large language models (LLMs) can help filter alarms, per-alarm prompting is often inaccurate and expensive. LLMs often misjudge such end alarms, and the repeated context across queries wastes many tokens. We shift LLM judgment from end alarms to intermediate facts (e.g., alias or flow edges), which are easier to validate. If a fact is judged false, all dependent facts and alarms can be pruned. We capture these dependencies in a derivation graph, enabling analyzer-agnostic pruning for any tool that exposes derivations. Under a token budget, we define the fact impact prioritization problem, which asks which facts to query first to maximize expected downstream pruning. We solve it with Bayesian program analysis by estimating each fact’s pruning impact from rule probabilities and fact posteriors. Building on these ideas, we present an LLM-based alarm resolution framework guided by Bayesian program analysis. It iteratively queries high-impact facts that LLMs can judge accurately, prunes downstream nodes when a fact is false, and feeds the judgments back to the Bayesian model as high-confidence evidence. We evaluate our approach on a Java datarace analysis and a C taint analysis, showing that it improves alarm-resolution quality while substantially reducing token consumption compared with both unfiltered static analysis and per-alarm LLM judging.