STV-FSANet: Track-Level Spatio-Temporal Verification for Fire and Smoke Alarm Validation in Video Surveillance
Deepak Ghimire, Donghoon Kim, Yeonho Jo, Eunhee Lee, Sunghwan Jeong, Byoungjun KimGenerating early and reliable fire/smoke alarms from real-world video surveillance remains challenging because visually ambiguous patterns such as sunlight, reflections, clouds, mist, steam, and illumination changes often trigger unstable frame-level false alarms. This paper presents the Spatio-Temporal Verification Network for Fire and Smoke Alarm Validation (STV-FSANet), a detect–track–verify framework that first localizes candidate fire/smoke regions, associates them into temporal tracks, and then verifies the observed sequence of each active track online as fire, smoke, or false fire/smoke. The verifier combines a primary appearance stream from cropped candidate regions with lightweight geometric cues derived from bounding-box position, scale, motion, and short-term fluctuation. A dual-branch GRU models long-term track history, while recent temporal pooling emphasizes newly observed evidence for streaming decisions. To support temporal learning, we construct the Fire–Smoke Alarm Verification (FSAV) Tracklet Dataset from 1347 source videos, yielding 58,733 parent tracks and 2.06 million annotated track frames. The best matched-context STV-FSANet achieves 97.09% test accuracy and 95.81% macro-F1, and both shorter-context models are within 0.5 percentage points of their final prefix metrics by 1.0 s. The results indicate that compact temporal evidence is more useful than simply accumulating longer 96-frame histories. The proposed model also rejects false fire/smoke tracklets with 90.2% recall, demonstrating the value of explicit hard-negative modeling, while the decoupled design allows the verifier to be reused with future detector backbones. TensorRT FP16 deployment reaches 109.55 frames/s on an RTX 4060 Laptop GPU and 39.51 frames/s on a Jetson AGX Orin DevKit.