Optimization of Spatial Downscaling Models for Satellite Imagery Based on Deep Learning and Generative Artificial Intelligence
Juan Valdés-Quintero, Rubén Darío Vásquez-Salazar, Juan Camilo Parra, César Olmos-Severiche, Andrés Gustavo Camargo-Perea, Cristian Alejandro Tibavija-Abril, Jean Pierre Díaz-PazSpatial downscaling of satellite imagery, the reconstruction of high-resolution outputs from coarser-resolution inputs, is a critical enabler of long-term land cover monitoring, yet the domain adaptation gap between natural-image super-resolution models and satellite sensor characteristics remains largely unaddressed. This study proposes a transfer learning framework for Landsat-to-Sentinel-2 spatial downscaling at a scale factor of ×3, systematically comparing shallow and deep fine-tuning strategies applied to two state-of-the-art architectures: SwinIR and ESRGAN. All models were trained and evaluated on a newly constructed dataset of 3500 paired patches assembled through an automated Google Earth Engine pipeline with rigorous cloud, water, and spectral variability filters spanning latitudes −30° to 30°. Performance is assessed using a hybrid evaluation framework combining pixel-wise metrics (MSE, PSNR, SSIM) with perceptual metrics (LPIPS, DISTS), and statistical significance is established through Friedman tests with Nemenyi post hoc analysis. Results demonstrate that off-the-shelf pretrained models fail to outperform bicubic interpolation on pixel-wise metrics, remaining statistically indistinguishable from it, thereby revealing a domain adaptation gap. Domain-specific fine-tuning completely reverses this degradation: ESRGAN with deep fine-tuning achieves the best performance across all five metrics simultaneously, reducing MSE by 48.7% and improving PSNR by 1.84 dB relative to the pretrained baseline, with all improvements statistically confirmed at p<0.0001. The findings establish that intermediate unfreezing of feature extraction blocks represents the optimal adaptation strategy, and that pixel-wise and perceptual metric families can diverge, making hybrid evaluation a methodological necessity rather than a redundancy.