DOI: 10.3390/s26165246 ISSN: 1424-8220

Component-Level Contributions of Retrieval-Augmented LLM Post-Processing in Streaming Anomaly Detection: A Matched-Operating-Point Case Audit

Changwon Baek

Retrieval-augmented large language model (LLM) post-processing reportedly improves anomaly triage over streaming industrial Internet of Things (IoT) sensor data, yet its LLM and retrieval contributions are rarely separated. We audit a build-verified pipeline (detector, LLM, retrieval, reranking), adding each stage at a matched detector recall of 0.9 under SHA-256-frozen preregistration, with bootstrap intervals and two generator sizes on 40 NAB and SKAB streams (19 industrial). Retrieval, the preregistered primary contrast, is null in seven of eight design configurations, but a preregistered grid of lower recall targets breaks that null on precision at both targets under stream-level resampling and in two of four cells when benchmark families are the unit. Across that grid, the LLM gain (72B F1 +0.0502) shrinks as candidate recall falls (+0.0227 at 0.7, undetectable at 0.5), which is a bound, not a dose–response curve. There, retrieval buys precision at a significant cost of recall. On the 19 industrial streams, the LLM gain is null at one label-free point and exactly zero at the other two, where the generator confirmed every candidate. Faithfulness, by local natural language inference, is low in every arm and not raised by retrieval. Three annotators labeling 133 claims (Fleiss’ κ = 0.555) placed the shortfall in the verdicts; none entailed by consensus. The protocol is released as a tested artifact.

More from our Archive