A Vision Transformer with Dynamic Masking and Cross-Modal Semantic Learning for Remote Sensing Scene Classification
Feng Ni, Yi Liu, Shibo Dai, Lei Chen, Changlei Feng, Fan Zhang, Xiang Wu, Yuming BoRemote sensing scene classification plays a vital role in various Earth observation applications. Although supervised learning remains the dominant paradigm, vast quantities of unlabeled imagery remain significantly underutilized. To leverage these unlabeled resources and enhance categorization accuracy, we propose a novel framework based on the Vision Transformer (ViT) that integrates a dynamic masking strategy with a cross-modal semantic learning mechanism. Specifically, a dynamic masking strategy guided by a smooth reconstruction loss is designed to learn robust feature representations from unlabeled samples prior to downstream fine-tuning. Furthermore, we incorporate cross-modal learning to enrich semantic information, thereby addressing the inherent supervisory limitations of conventional one-hot labels. Comprehensive experiments demonstrate that the proposed method significantly improves classification accuracy while maintaining high pre-training efficiency and strong generalization capabilities.