DOI: 10.3390/geomatics6050111 ISSN: 2673-7418

A Comparison of Machine Learning and Deep Learning Architectures with Explainable AI for Mapping Land Cover Using Sentinel-2 Satellite Data

Muhammad Salman Younas, Muhammad Usman, Muhammad Ali, Roman Shults, Muhammad Hashir, Muhammad Bilal, Zubair Nawaz, Md Masudur Rahman

Rapid and unplanned urbanization in South Asian cities such as Lahore, Pakistan has produced fragmented urban-agricultural landscapes in which spectral signatures of various land cover and land use (LCLU) classes overlap, degrading the performance of conventional pixel-based machine learning classifiers. This study presents an explainability-driven comparison of a pixel-based machine learning classifier and five deep learning semantic segmentation architectures for LCLU mapping using Sentinel-2 imagery in Lahore District. Extreme Gradient Boosting (XGB) was compared against U-Net and U-Net++ (dense convolutional networks), DeepLabV3+ (atrous convolution), CNN-BiLSTM (recurrent–convolutional hybrid), and SegFormer (vision transformer). Each model was evaluated across five combinations of input features using different spectral bands and indices from Sentinel-2 satellite imagery, and each deep learning architecture was additionally trained under two initialization strategies: ImageNet-pretrained weights adapted through a non-linear multispectral stem and training from scratch, producing 55 experimental scenarios evaluated on a spatially structured (radial) hold-out subset with a 500 m buffered exclusion zone to mitigate spatial autocorrelation. U-Net++ showed the best overall performance (overall accuracy 94.28%, macro F1 91.86, Cohen’s Kappa 91.87, mean IoU 85.19), exceeding the best XGB configuration by 24.82% points in mIoU. Incorporating the complete set of ten spectral bands from Sentinel-2 as input features improved the U-Net++ mIoU from 79.74 to 85.12 over the RGB+NIR combination, with the Red-Edge and short-wave infrared (SWIR) bands being particularly helpful for discrimination of vegetation and water, respectively. The use of indices added little benefit for the deep networks but remained important for machine learning models. The value of ImageNet pretraining was strongly architecture-dependent: it improved the weaker architectures substantially (DeepLabV3+ and CNN-BiLSTM with +2.82 and +2.40 average mIoU points, respectively) but was negligible for the U-Net family and consistently detrimental for SegFormer. SHAP and Grad-CAM analyses indicated that the machine learning baseline relied on isolated spectral channels, whereas the convolutional networks exploited spatial structure (edges, texture, and morphology) to resolve spectral ambiguity. These results indicate that, for fragmented urban-agricultural landscapes with limited labeled data in our study area, convolutional encoder–decoder architecture remains the most reliable choice, and that generic visual pretraining is beneficial only where an architecture lacks strong spatial priors of its own.