DOI: 10.3390/fire9100427 ISSN: 2571-6255

MFFLNet: Multi-Scale Fire Feature Learning for Fire Video Recognition

Shanzheng Yang, Yun Yi

Accurate and timely recognition of flame and smoke is critical for fire warning systems to mitigate casualties and property damage. Most existing fire datasets are image-based and thus fail to capture the dynamic spatiotemporal information of fire. Furthermore, the lack of large-scale fire video datasets poses a significant challenge to the training of neural networks. To address these limitations, we developed the Flame-Smoke Video Recognition (FSVR) dataset, a large-scale collection of 30,000 video clips that significantly exceeds the scale of prior datasets in the same domain. Since flame and smoke exhibit multiscale characteristics, small-scale targets are often obscured by complex backgrounds. Existing methods lack robust multiscale feature learning capabilities, which limits their ability to achieve precise fire recognition under complex conditions. To address this issue, we proposed the Multi-scale Fire Feature Learning Network (MFFLNet), which integrates multiple Multi-scale Fire Feature Learning (MFFL) blocks into a Transformer backbone. Each MFFL block comprises two key components, i.e., the multiscale fire Conv3D layer and the fire spatiotemporal feature learning layer. The experimental results obtained from the FSVR and LFVR datasets demonstrate that MFFLNet surpasses the baseline model and other comparative methods. When the backbone network is initialized with pre-trained weights from the Kinetics-710 dataset, MFFLNet attained an accuracy of 79.54% and a macro-F1 score of 79.22% on the FSVR dataset, while achieving an accuracy of 95.64% and an F1 score of 95.46% on the LFVR dataset.