Depth Adaption SegNet for RGB-T Segmentation
Shaochuan Zhao, Chi Zhang, Hancheng Zhu, Bing Liu, Yong ZhouRGBT segmentation is a challenging task in the area of computer vision. Current advanced networks for RGBT segmentation focus on extracting deeper discriminative features from RGB and thermal images to provide richer semantic information for the fusion features to the decoder. However, excessively mining deeper semantic features only makes the model redundant. Simultaneously, lacking shallow spatial features leads to difficulties in guaranteeing accurate localization of targets. We believe that the features provided by images can be categorized into three types: edge, patch, and semantics. Only by synchronously taking into account the extraction of all three types of features can models achieve accurate classification on the basis of precise localization. Therefore, we propose Depth Adaption SegNet for RGB-T Segmentation (DASNet). According to the characteristics of the three types of features, we extract semantics, patch, and edge features from the deep, middle, and shallow stages respectively. We specifically design the cross-attention semantics module, patch activation module, and edge enhancement module to perform feature extraction. In addition, in order to efficiently fuse features from different categories, we design a deep-emphasis fusion module to fuse the output features of the modules. Compared to advanced methods, qualitative and quantitative experiments show that DASNet exhibits state-of-the-art performance on the CNN-based RGBT segmentation task.