Temporal-Variation-Resistant Bidirectional Convolution-Transformer GAN for Remote Sensing Image Spatiotemporal Fusion
Yuanyuan Wu, Linjie Fu, Xinying Zhong, Yuxuan Qiu, Cong LinSingle-source remote sensing image (RSI) cannot simultaneously meet high-spatial and high-temporal resolution requirements, failing to provide decision-makers with timely and accurate monitoring data. Spatiotemporal fusion (STF) of multi-source RSIs represents an efficient and convenient means of producing land-cover observations with high-temporal and high-spatial resolutions. However, current STF approaches still suffer from severe prediction distortion under abrupt changes, long-interval temporal variations, and land-cover type transitions, as well as poor robustness against disturbances in prior data. To address these challenges, a temporal-variation-resistant bidirectional convolution-Transformer generative adversarial network (TRB-GAN) for RSI STF, which comprises a temporal-variation-resistant bidirectional convolution-Transformer generator (TRBG) and a multiresolution input convolution-Transformer discriminator (MICTD), is devised to improve the robustness in predicting time-varying information and enhance STF capability. First, the TRBG designs a temporal-variation-resistant bidirectional encoder to capture prior information and arbitrary time-varying local–global features, enhancing prediction robustness and representation capability for time-varying information. Second, the TRBG designs a dual-guided triple-attention fusion decoder (DTAFD), incorporating dual-guided cross convolution-attention fusion and decision attention fusion. DTAFD dynamically calculates correlations among spectral, spatial, and time-varying information to aggregate heterogeneous features and adaptively performs stepwise weighting and integration, effectively mitigating the adverse impacts from heterogeneous imaging mechanisms and significant resolution gaps. Finally, MICTD and deep supervision enable adversarial learning of local–global structures and spectra across resolutions, providing feedback to the TRBG for producing finer images. Ablation and comparative experiments demonstrate the TRB-GAN achieves superior STF performance and stronger robustness to time-varying disturbances for the widely used CIA and LGC datasets.