CMST-Net: Cross-Modal Interaction and Spatio-Temporal Feature Enhancement Method for Continuous Sign Language Recognition
Qiuhong Tian, Zhengzheng Li, Hanbo Zhang, Shiwei Ge, Jing HuangIn continuous sign language recognition (CSLR), existing methods predominantly adopt frame-wise feature extraction, neglecting temporal continuity and motion trajectory modeling, thereby struggling to capture complete spatio-temporal dynamics. Meanwhile, global cross-modal attention approaches typically directly model interactions between text and the entire video sequence but lack structural constraints, making them susceptible to interference from redundant frames, which leads to attention distribution dilution and undermines fine-grained motion alignment capability. To address these issues, this paper proposes CMST-Net, a Cross-modal Interaction and Spatio-temporal Feature Enhancement Method for Continuous Sign Language Recognition. CMST-Net comprises two principal modules: the Local–Global Cross-modal Fusion Module (LGCFM) and the Spatio-Temporal Feature Enhancement Module (STFEM). LGCFM introduces a synergistic modeling mechanism that combines local sliding-window attention with global attention, achieving a unified fusion of structured local alignment and global semantic modeling. STFEM incorporates multi-scale spatial dilated convolution and coordinate attention to extract fine-grained spatial features while leveraging channel partitioning and a hierarchical residual structure to enhance long-range temporal modeling capability; the two modules collaboratively yield high-quality spatio-temporal feature representations. Experiments on three public benchmark datasets (PHOENIX2014, PHOENIX2014-T, and CSL-Daily) demonstrate that CMST-Net can effectively improve continuous sign language recognition performance, achieving state-of-the-art performance on the PHOENIX2014 and CSL-Daily datasets and competitive results on the PHOENIX2014-T dataset.