DOI: 10.3390/rs18162706 ISSN: 2072-4292

EdgeNeXt-Attn: A Lightweight Attention-Enhanced Deep Learning Framework for Fire Detection in Remote Sensing Imagery

Hikmat Yar, Nehad Ali Shah, Weiwei Jiang, Norah Saleh Alghamdi, Heung Soo Kim

Wildfires are a major environmental hazard with severe consequences for ecosystems, air quality, infrastructure, and public safety. The rising incidence and severity of wildfire events worldwide have increased the need for reliable early detection and monitoring systems. Remote sensing technologies, such as satellite and unmanned aerial vehicle (UAV) imagery, along with ground-based Closed-Circuit Television (CCTV) cameras, provide valuable geospatial data for large-scale wildfire monitoring. Recent advances in deep learning, particularly Convolutional Neural Networks (CNNs) and Transformer-based architectures, have significantly improved the accuracy of wildfire detection systems. Despite these advances, balancing local feature representation with global contextual modeling remains challenging. CNNs effectively capture local spatial features but have limited receptive fields, whereas Vision Transformers (ViTs) model long-range dependencies but often overlook fine-grained local details and require substantial computational resources. Consequently, accurately detecting small, occluded, and visually ambiguous fire regions remains difficult, particularly for real-time deployment on resource-constrained edge devices. To address these challenges, this study proposes EdgeNeXt-Attn, an enhanced EdgeNeXt-based framework that effectively integrates local feature learning and global contextual modeling through channel and spatial attention mechanisms. The proposed model improves the detection of small, occluded, and visually ambiguous fire regions while maintaining the computational efficiency required for real-time edge deployment. The proposed framework is evaluated on four multi-platform benchmarks spanning ground-based CCTV (DFAN, Complex-Fire), aerial drone (FLAME), and mixed drone–satellite (ADSF) imagery, achieving 92.09%, 95.16%, 96.65%, and 87.81% accuracy, respectively, and outperforming recent state-of-the-art baselines. With only 5.3M parameters, the model achieves real-time inference at 85.9, 27.3, and 8.4 FPS on GPU, CPU, and Raspberry Pi, respectively. Furthermore, ablation studies and Grad-CAM analysis validate its effectiveness and accurate fire localization. These results demonstrate an accurate and computationally efficient framework for real-time wildfire monitoring using multi-platform remote sensing and ground-based imaging systems.

More from our Archive