DOI: 10.3390/rs18193363 ISSN: 2072-4292

Dual-Domain Triple-View Contrastive Representation Learning for Few-Shot Hyperspectral Image Classification

Rujun Zhang, Wenming Cao, Qifan Liu

Few-shot hyperspectral image classification (FS-HSIC) remains highly challenging because the limited number of labeled samples not only makes traditional supervised learning models prone to overfitting but also restricts their ability to learn generalizable spatial–spectral representations. Moreover, existing methods are typically confined to a single Euclidean space and may therefore fail to capture complementary geometric structural information. Traditional contrastive learning based on two augmented views may also discard the information in the original samples that is essential for distinguishing different classes. To address these issues, we propose a dual-domain representation learning framework that jointly leverages complementary representations from the Euclidean and geometric algebra (GA) domains for FS-HSIC. In the Euclidean domain, we develop a multi-scale spatial–spectral residual Transformer (MSSRT) network that aggregates multi-scale long-range contextual information by progressively enlarging the attention windows and introducing cross-stage residual connections. A spatial–spectral relation modeling (SSRM) module further learns spatial dependencies and spectral channel correlations in parallel. In the GA domain, a spatial–spectral geometric algebra Transformer (SS-GATr) encodes hyperspectral features as multivectors and captures geometric information complementary to that represented in the Euclidean representations. During contrastive learning, an additional original-view branch is incorporated alongside two augmented-view branches, enabling the network to preserve fine-grained information while learning representations that are invariant to data augmentation. Furthermore, we propose a bidirectional guide-conditioned domain interaction (BGDI) module that selectively transfers information between the two domains based on inter-branch feature correlations, thereby facilitating the bidirectional exchange of complementary information while suppressing interference from irrelevant features. During few-shot finetuning and classification, the predictions of the two branches are weighted using a controllable fusion coefficient to balance the contributions of the two complementary representations when only limited labeled samples are available. Experiments are conducted on the Indian Pines, Pavia University, and Houston datasets using only five labeled samples per class. Compared with the state-of-the-art methods, the proposed method achieves overall accuracies (OAs) of 77.15%, 84.63%, and 83.46% on the three datasets, respectively, demonstrating the strong and competitive performance.