Modality Effectiveness in Cross-Subject Sign Language Recognition: A Skeleton Perspective
Sijie Miao, Jiaxin He, Lu ZengAccessibility-oriented sign language recognition plays an important role in intelligent human–computer interaction, but cross-subject recognition remains challenging because test signers are unseen during training. Variations in appearance, motion style, and modality quality among different signers introduce significant distribution differences between training and testing samples, limiting model generalization. Although multimodal approaches commonly assume that RGB, skeleton, and depth modalities provide complementary information, their effectiveness under cross-subject recognition remains insufficiently explored. This study investigates modality effectiveness in SLR500 word-level sign language recognition under a cross-subject protocol. A unified evaluation framework is established to compare RGB-only, skeleton-only, RGB–skeleton fusion, and RGB–skeleton–depth fusion under consistent data splits, preprocessing procedures, training settings, and evaluation metrics. Sample-level prediction transitions are further analyzed to examine how fusion changes recognition outcomes relative to the skeleton-based reference model. The results show that the skeleton modality provides more reliable recognition performance under the evaluated cross-subject setting, while the tested fusion strategies do not consistently improve over the skeleton-only configuration. Based on the modality-effectiveness analysis, a skeleton-based reference model is established using a Spatial–Temporal Graph Convolutional Network (ST-GCN)-based spatio-temporal encoder, learned confidence-weighted temporal pooling, and a lightweight classifier. Multi-seed experiments further verify the stability of the model, which achieves 86.37% Top-1 accuracy and 96.86% Top-5 accuracy on the SLR500 cross-subject test set. These results emphasize the importance of evaluating modality reliability before designing complex fusion strategies for cross-subject sign language recognition.