Forecasting Deployment‐Time Blind Spots in Language‐Model Intrusion Detection Agents: A Data‐Geometric Coverage Statistic and the Relay Condition That Carries It to the Agent
Qianli Di, Keju Du, Gang Chen, Hai DengABSTRACT
An operator deploying an intrusion detector on a new attack surface cannot say which attack families it will silently miss. For tree‐ensemble detectors, that can be forecast before deployment from labelled traffic, with confirmation on one of the two benchmarks we test. For a candidate family, a nearest‐neighbour coverage statistic computed from labelled known traffic and a small sample of that family uses no model output and nothing about the detector. On the 11 CIC‐IDS2017 leave‐one‐attack‐out folds the forecast survives rank correlation, bootstrap confidence intervals and multiplicity correction for both tree‐ensemble families (random forest , 0.87 across reference‐set draws, CI [0.68, 0.99]); on UNSW‐NB15 it holds directionally only. Ten labelled samples suffice ( against at 200). Reaching the deployed language‐model agent is what harness engineering contributes: SecHarness, a five‐component execution‐layer specification, constrains a 3B‐parameter agent into a relay whose blind spots coincide with its tool's. We report the null in full: against the same random‐forest verdict written into the prompt, the harness changes 1 decision in 200 at latency, and fine‐tuning adds neither accuracy nor speed on matched hardware, while the fine‐tuned 3B agent fabricates tool citations once its tool is withdrawn. It keeps rather than creates what the tool supplies: catching every error costs 40.5% of the alert stream under the harness against the detector's own 43.5%, a six‐record gap we report as an artefact. Knowing where an agent will be blind should precede optimising the model inside.