DOI: 10.3390/electronics15194470 ISSN: 2079-9292

A Hybrid Deep Learning Approach for Early Detection of Wildfires Using Transformer Models and Feature Fusion Techniques

Abdullah Şener, Burhan Ergen, Kubilay Demir, Serdar Ekinci, Vedat Tümen

Natural disasters, particularly wildfires, pose a serious threat to ecosystems, human lives, and infrastructure, making rapid and accurate fire detection essential for effective disaster management. Despite recent advances in deep learning-based wildfire detection, conventional CNN-based approaches are inherently limited in modeling long-range spatial dependencies, while relying on a single deep architecture may provide an incomplete representation of complex and multi-scale fire patterns. Furthermore, high-dimensional deep features may contain redundant information, increasing computational complexity and limiting the efficiency of real-time monitoring systems. Therefore, there remains a need for an efficient framework that can exploit complementary global representations from multiple transformer architectures while retaining only the most discriminative features. To investigate this application-specific problem, this study evaluates a hybrid framework that integrates three heterogeneous Vision Transformer architectures—DeiT3, MaxViT, and Swin—with feature-level fusion, mRMR-based feature selection, and SVM classification for binary wildfire image classification. The publicly available FlameVision dataset, consisting of aerial images categorized into fire and non-fire classes, was used for evaluation. Deep feature representations extracted from the three transformer architectures were systematically fused to exploit their complementary characteristics. Minimum Redundancy Maximum Relevance (mRMR) feature selection was subsequently employed to reduce feature redundancy and retain the most discriminative information, followed by classification using Support Vector Machines (SVM). The experimental results demonstrated an overall classification accuracy of up to 100% on the internal hold-out test partition and in five-fold cross-validation within the FlameVision dataset. Because both evaluation procedures were conducted using images originating from the same source dataset, these results should be interpreted as within-dataset classification performance rather than evidence of external generalization. Validation on independently collected wildfire datasets is therefore required to determine the robustness of the proposed framework under different geographical, sensor, environmental, and acquisition conditions. The proposed framework demonstrates that complementary transformer-derived representations combined with feature selection can provide highly discriminative wildfire representations while substantially reducing the number of features required for classification. These findings demonstrate that the proposed approach can retain high classification performance using a substantially more compact feature representation. However, the resulting reduction in feature dimensionality should not be interpreted as evidence of improved runtime efficiency, since computational time before and after feature selection was not experimentally measured. However, real-time performance is not claimed in the present study because FPS, end-to-end latency, runtime memory consumption, and deployment on continuous UAV or satellite streams were not experimentally evaluated. Such operational evaluation is left for future work.