Protocol-Conditioned Feature Analysis, Multi-Dataset Intrusion Detection, and Context-Aware Alert Prioritization for IoT Network Security
Abdullah Abbasi, Dil Nawaz Hakro, Akhtar Hussain, Osama Alrahbi, Mohammed Izaan Kari, Suad Mohammed Al Qassabi, Muhammad Hafidz Fazli Bin Md FauadiInternet of Things (IoT) malware and intrusions generate network-observable flow patterns, but high benchmark accuracy does not by itself establish transfer to new environments. This study presents Protocol-conditioned Analysis with Behavioral Threat Identification (PA-BTI) as an evidence-bounded framework combining dataset-specific Random Forest/XGBoost detection, protocol-conditioned direction-invariant CorrAUC (dCorrAUC) diagnostics, a River streaming diagnostic, and context-aware alert prioritization. Across the primary evaluations, selected operating points reached 99.77% accuracy with 0.48% false-positive rate (FPR) on ToN-IoT and 99.44% accuracy with 0.08% FPR on a balanced Bot-IoT holdout, whereas the official NSL-KDD split reached 79.55% accuracy and 66.20% attack recall. To address testbed-artifact sensitivity, a new official UNSW-NB15 train/test ablation obtained 87.67% accuracy, 98.59% attack recall, and 25.70% FPR with all features; removing TTL/state variables changed accuracy to 87.22%, and additionally removing the ct_* windowed features changed it to 86.31%, showing no collapse but a persistently high official-split FPR. A new integrated common-flow replay propagated real UNSW detector outputs and training-only protocol evidence through the contextual layer using explicitly synthetic CTI/asset context. On a disjoint 41,166-record evaluation half, calibrated fusion did not improve context-graded nDCG@100 over confidence alone (0.610 versus 0.627), demonstrating component interaction while bounding any claim of prioritization benefit. The River diagnostic also exposed a minority-class failure mode (91.56% accuracy but attack-class F1 = 0.003). PA-BTI is therefore presented as a reproducible, bounded evaluation framework rather than evidence of cross-dataset generalization or robust portability.