DOI: 10.3390/rs18193291 ISSN: 2072-4292

RGB-to-Multispectral Reconstruction with Spectral Attention and Non-Adversarial Perceptual Loss

Hanyu Liu, Hongrui Xing, Lipeng Pan

Multispectral remote sensing provides rich spectral information for land monitoring and environmental analysis, but sensors like Sentinel-2 are often constrained by acquisition limitations and cloud cover, while widely accessible RGB imagery lacks detailed spectral bands. Reconstructing multispectral bands from RGB is, however, a severely ill-posed inverse problem: the red-edge, near-infrared (NIR), and short-wave infrared (SWIR) bands share no direct pixel-wise correspondence with the three visible channels, and under cloud-contaminated acquisitions part of the spectral information is entirely missing. Existing deep models mostly rely on adversarial training, which is unstable and prone to introducing spectral distortion, whereas models trained with pixel-wise losses alone tend to over-smooth fine details. This paper proposes a reconstruction framework to generate ten selected Sentinel-2 bands (B2–B8, B8A, B11–B12) from RGB images. The generator uses a residual U-Net with a spectral attention module at the bottleneck, trained with a composite loss combining pixel-wise L1 and perceptual LPIPS losses, balancing spectral fidelity and structural realism. On the fMoW-Sentinel test set, the L1+LPIPS model achieves a PSNR of 20.28 dB, an SSIM of 0.844, an LPIPS of 0.326, and an SAM of 6.80, outperforming more complex GAN-based variants on all four metrics; adding LPIPS alone improves PSNR by 0.58 dB over the L1-only baseline, and the spectral attention module yields a further gain. Without any fine-tuning, it retains PSNRs of 19.82 and 19.68 dB on EuroSAT and BigEarthNet-S2, respectively. Architecture ablation confirms that the spectral attention module provides consistent improvements, and cross-dataset generalization experiments demonstrate transferability. Downstream NDVI evaluation validates the practical utility of the reconstructed multispectral imagery. The framework offers a practical solution for spectral enhancement, with modular architecture enabling future extension.