DOI: 10.3390/app16157691 ISSN: 2076-3417

GLA-DesnowNet: A Lightweight Hybrid CNN–Transformer Architecture for Image Snow Removal

Habibulloyev Fakhriddin Abduhalim Ugli, Mst Farjana Aktar, Unal Aras, Tulkinov Bakhromjon Nusratjon Ugli, Jee Youl Ryu, Tahesin Samira Delwar

Single-image snow removal remains a challenging, ill-posed inverse problem in computer vision due to the highly variable appearance of snow degradation. Existing CNN-based methods are limited by local receptive fields and cannot model globally distributed snow patterns, while Transformer-based methods achieve strong performance at a prohibitive computational cost. To address both limitations, GLA-DesnowNet is proposed as a lightweight hybrid architecture that introduces a novel Gated Transformer (GT) module, a gated global attention mechanism designed specifically for physics-driven image desnowing. The Gated Transformer applies global multi-head self-attention with a learned convolutional gate, enabling long-range spatial dependencies to be captured while remaining computationally efficient. An adaptive three-channel snow-mask generation strategy is introduced to automatically produce pixel-level supervision from paired images without manual annotation. A compound loss combining L1 reconstruction with multi-scale pyramid loss is employed to ensure consistent restoration across spatial scales. Evaluated on the Snow100K benchmark, the proposed method achieves 29.30 dB PSNR and 0.93 SSIM in only 160 training epochs, with 0.53 M parameters and 4.44 GFLOPs per inference—significantly fewer than existing state-of-the-art (SOTA) methods—while maintaining competitive restoration quality. On the Comprehensive Snow Dataset (CSD), the identical model achieves a 27.79 dB PSNR and 0.90 SSIM. These results confirm a strong efficiency–accuracy trade-off and cross-dataset generalization suited for resource-constrained and real-time deployment.

More from our Archive