Explainable and Analyst-Driven Random Forest for Intrusion Detection
Saloua Bellouch, Mostapha Zbakh, Siham Aouad, An BraekenRandom Forest and other tree-ensemble classifiers achieve high accuracy in network intrusion detection; however, their aggregate decision logic prevents analysts from auditing or deploying individual predictions as operational rules. Post hoc explanation methods introduce latencies incompatible with security operation center (SOC) requirements and produce conditions unsuitable for firewall configuration. Among the systems reviewed in this study, none unifies intrinsic explanation, rule deployment, ATT&CK attribution, cross-dataset validation, and adaptive feedback in one pipeline. This work presents a depth-limited Random Forest with deterministic, per-instance explanations at a fraction of gradient-based attribution latency. Complementary mechanisms generate analyst-deployable rule specifications, technique-level adversary attribution, and a feedback protocol that models label noise, missed reviews, and bounded correction budget. Evaluated on a large, multi-category network-traffic benchmark, the system attains high detection accuracy (macro recall 0.86, driven substantially by the majority normal-traffic class at 68% of flows) while sustaining throughput beyond SOC requirements; a stealthy reconnaissance-and-exploitation category remains markedly harder to detect under this class imbalance. Cross-dataset evaluation on a more recent benchmark attains strong performance after limited target-domain retraining. The adaptive feedback protocol yields a statistically significant false-positive reduction over repeated simulated reviews, requiring only modest weekly analyst effort. Together, these capabilities enable auditable and SOC-integrable detection pipelines.