C2DSSL: Context-Consistency Enhanced Collaborative Self-Supervised Learning for Remote Sensing Image Understanding
Wu Wen, Jinghui Luo, Kailun Qiu, Zhong Xiao, Gen LaiSelf-Supervised Learning (SSL) has attracted increasing attention in remote sensing image understanding because it can learn transferable representations from unlabeled images. However, two issues remain insufficiently examined in collaborative SSL for remote sensing. First, when high-ratio masking removes entire small objects or structurally informative regions, the remaining visible patches may provide insufficient evidence for semantically coherent reconstruction. Second, heterogeneous self-supervised objectives may exhibit different loss scales, gradient magnitudes, and convergence behaviors such that fixed coefficients can produce uneven branch contributions during training. To address these issues, this paper proposes Context-Consistency Enhanced Collaborative Self-Supervised Learning (C2DSSL) for remote sensing image understanding. C2DSSL introduces a teacher–student context-consistency constraint, in which multi-scale reconstruction features from the complete teacher observation serve as contextual targets for the student network when reconstructing the corresponding masked observation. In addition, gradient-sensitive dynamic weighting uses temporally smoothed loss-gradient magnitudes as empirical signals to adjust the relative contributions of heterogeneous self-supervised objectives. Under the evaluated settings, adding the context-consistency constraint improves KNN representation evaluation, UCMerced classification, Potsdam semantic segmentation, and DOTA oriented object detection, while maintaining the Baseline performance on the Million-AID subset. Dynamic weighting shows task-dependent effects, including a performance gain on the Million-AID subset classification task when combined with CCL.