DOI: 10.3390/rs18193281 ISSN: 2072-4292

CoGLoR: Collaborative Global–Local Representation Learning for Remote Sensing Scene Classification

Jingyu Xu, Ruyue Feng, Ziyou Guo, Tieru Wu, Huilai Li

Self-supervised learning provides an effective way to reduce the dependence on manual annotations in remote sensing scene classification. However, existing self-supervised methods often emphasize either global semantic invariance or local structural modeling, while the interaction between global scene context and local patch-level representations remains insufficiently explored. This limitation is important for remote sensing images, whose scene categories are usually determined by both holistic spatial layouts and fine-grained local objects. To address this issue, we propose CoGLoR, a collaborative global–local representation learning framework for self-supervised remote sensing scene classification. The proposed framework jointly integrates global contrastive learning, masked reconstruction, and Contrastive Patch-Level Attention Matching within a unified training pipeline. The contrastive branch learns scene-level semantic consistency across augmented views, while the reconstruction branch encourages the encoder to preserve local structural details by reconstructing masked patches. C-PAM replaces fixed spatial indices with a hard top-K, multi-positive feature neighborhood and contrasts the selected neighbors against batch-level candidates. The selected features are treated as learned neighbors rather than verified semantic correspondences. Experiments on AID, UCMerced Land Use, and NWPU-RESISC45 show that CoGLoR improves over representative contrastive learning, masked modeling, and hybrid self-supervised baselines under linear probing and low-label evaluation protocols. The matched ResNet-50 comparison supports the method-level evaluation; the retrieval figure illustrates cosine-ranking behavior; the ViT-S/16 experiment tests backbone portability; and the resource table quantifies pretraining overhead. Externally pretrained methods are included only as contextual references.