DOI: 10.3390/app16199640 ISSN: 2076-3417

Multi-Scene Continuous Sign Language Recognition Based on Temporal–Frequency Enhancement and Discrete Cosine Transform Linear Attention

Xiangyang Sun, Chuhan Wang, Zihan Cai, Wenjun Zhang, Binggao He

To address the adverse effects of complex acquisition conditions on continuous sign language recognition (CSLR) performance and the increasing computational cost of standard self-attention with increasing video sequence length, this paper proposes a continuous sign language recognition method based on temporal–frequency enhancement and DCT linear attention. First, a multi-scene continuous sign language dataset was constructed from the recordings of nine signers. The dataset contains 3502 video clips, 1208 lexical items, and nine acquisition scenarios, covering indoor and outdoor environments, strong and weak illumination, and static and dynamic backgrounds. These acquisition scenarios were designed to introduce diverse visual conditions during data collection. Because the available evaluation uses a random video-level split without scene-disjoint grouping, the results describe performance under the recorded mixed conditions and do not establish generalization to unseen signers or unseen acquisition scenarios. Second, using a Video Swin Transformer and spatial global average pooling as the feature-extraction front end, a temporal–frequency-enhanced teacher model was developed through a temporal branch, a short-time Fourier transform (STFT) frequency branch, and gated fusion, thereby jointly exploiting temporal and local frequency information. On this basis, a lightweight student configuration was developed using DCT-kernelized linear attention in the temporal-modeling pathway. Response-level knowledge distillation was further introduced to mitigate the recognition-performance loss associated with lightweight linearized modeling. Experimental results show that the teacher model achieves word error rates (WERs) of 19.8%, 23.1%, and 37.2% on PHOENIX14, CSL-Daily, and the self-constructed multi-scene dataset, respectively. After knowledge distillation, the student model achieves WERs of 21.2%, 24.1%, and 38.8% on the three datasets, with gaps of 1.4, 1.0, and 1.6 percentage points relative to the teacher model, respectively. At the complete-configuration level, the reported FLOPs decrease from 12.5 G for the teacher model to 4.2 G for the student model. On an NVIDIA RTX 3090 Ti GPU with a batch size of 1 and a 200-frame input, the single-sample inference latency decreases from 48.5 ms to 36.8 ms. These results indicate that, under the specified experimental conditions, the lightweight student configuration achieves lower computational cost and inference latency at the expense of only a limited loss in recognition performance.