DOI: 10.3390/electronics15163581 ISSN: 2079-9292

Action Recognition Method Based on Multi-Scale Dilated Feature Fusion and Decoupled Spatiotemporal Attention Pooling

Hanbo Zhang, Jing Huang

Video action recognition requires the joint modeling of spatial appearance information and temporal dynamics. However, existing efficient action recognition methods based on two-dimensional convolution still have limitations in representing multi-scale spatial cues and aggregating key spatiotemporal information. To address these issues, this paper proposes an action recognition network based on Multi-Scale Dilated Feature Fusion and Decoupled Spatiotemporal Attention Pooling, termed MDSTA-Net. The proposed method adopts TSM as the basic temporal modeling framework and ResNet-50 as the backbone network. First, a Multi-Scale Dilated Feature Fusion module (MSDF) is designed to construct continuous multi-scale receptive fields through parallel convolutional branches with different dilation rates. An adaptive branch aggregation mechanism is further introduced to dynamically fuse responses at different scales, thereby enhancing the representation of both local details and broader contextual information. Second, a Decoupled Spatiotemporal Attention Pooling module (DSTAP) is proposed to model key action frames along the temporal dimension and salient discriminative regions along the spatial dimension. A residual pooling path is also incorporated to preserve global semantic information, improving the discriminative capability of video-level action representations. Experimental results on three public datasets, namely Something-Something V2, Kinetics-400, and HMDB51, demonstrate that MDSTA-Net achieves favorable recognition performance compared with several representative methods. Ablation studies further verify the effectiveness of MSDF and DSTAP, indicating that multi-scale spatial feature enhancement and key spatiotemporal information aggregation can effectively improve action recognition performance.

More from our Archive