DOI: 10.3390/a19090804 ISSN: 1999-4893

Algorithm Design and Analysis of a Blockchain-Enabled Reinforcement and Active Learning Framework for Robust IoT Intrusion Detection

Nadeem Javaid

The fast development of Internet of Things (IoT) gadgets has brought both a new level of connectivity and automation, at the same time making the networks more vulnerable to more sophisticated cyber attacks. Existing intrusion detection systems have a number of fundamental limitations, such as extreme class imbalance in network traffic data, inadequate feature interaction modeling, inability to train on dynamic attack patterns, heavy reliance on fully labeled data, weak evaluation capabilities, and limited interpretability. To deal with these issues, this research employs a data balancing mechanism with an autoencoder to reduce class imbalance by learning meaningful latent representations of the minority attack classes. A Parallel Hybrid single-step BiLSTM-FCNN (PH-BiLSTM-FCNN) model is then proposed as a single network to learn contextual dependencies and nonlinear discriminative features simultaneously along parallel learning pathways. To further enhance adaptability and convergence stability, a Q-learning Optimized PH-BiLSTM-FCNN (Q-PH-BiLSTM-FCNN) is introduced, enabling dynamic optimization of training behavior based on feedback-driven interactions. In addition, an Entropy-based Active Learning on PH-BiLSTM-FCNN (EAL-PH-BiLSTM-FCNN) is developed to significantly reduce labeling requirements by selectively querying the most informative samples. Furthermore, a blockchain layer with smart contracts is embedded consistently in all of the proposed models to ensure tamper-resistant logging, transparent validation, and reliable recording of training and evaluation results. From an algorithm design perspective, the proposed framework is structured as a set of verifiable procedures for PH-BiLSTM-FCNN training, Q-learning optimization, and entropy-based active sample selection. Its algorithmic behavior and robustness are evaluated through execution-time analysis, 10-fold cross-validation, and permutation feature importance, ensuring both predictive reliability and interpretability. Experimental results demonstrate that the proposed PH-BiLSTM-FCNN, Q-PH-BiLSTM-FCNN, and EAL-PH-BiLSTM-FCNN models consistently outperform state-of-the-art FCNN, Bi-LSTM, GRU, LSTM, LR models, achieving improvements of 4.80%, 13.52%, and 6.85% in accuracy, 4.81%, 13.53%, 6.86% in recall, and 5.18%, 18.44%, and 8.52% in precision recall-area under the curve, respectively. These results affirm that the proposed framework provides improved detection performance, statistical reliability, and interpretability while addressing practical deployment limitations in IoT security settings.