A Validation-Controlled Label-Efficient Framework for Coastal Wetland Habitat Mapping Using Multi-Season Sentinel-1 and Sentinel-2 Data
Marwa Zerrouk, Siham Fellahi, Asmaa Moussaoui, Imane Sebari, Kenza AitelkadiReliable coastal wetland habitat mapping is often constrained by the scarcity and the cost of reliable reference data, especially in data-limited coastal environments. We propose a validation-controlled, label-efficient framework pairing multi-season Sentinel-1 and Sentinel-2 predictors with a CatBoost teacher and a lightweight MLP student. A candidate is pseudo-labeled only when both separately calibrated models agree and exceed class-specific thresholds; accepted labels are class-balanced and down-weighted. The framework was evaluated at the Sidi Moussa–Oualidia wetland complex and Merja Zerga lagoon in Morocco. At Sidi Moussa–Oualidia, 62 configurations were compared through nested polygon-grouped validation and then frozen before a five-seed held-out evaluation. The supervised MLP and Agreement-augmented MLP achieved mean Macro-F1 values of 0.9518±0.0044 and 0.9509±0.0062, indicating that augmentation did not materially change the already strong full-data baseline. Under a stricter budget of 30 training and 20 validation observations per class, Agreement yielded a mean Macro-F1 of 0.9092±0.0102 compared with 0.9023±0.0093 for the supervised baseline and produced pseudo-labels in all five seeds. A spatial-range sensitivity analysis further showed that both models retained Macro-F1 values of 0.9391 and 0.9403 for test observations located beyond the largest estimated within-class autocorrelation range. At Merja Zerga, the native six-class supervised MLP achieved 0.9456±0.0050, compared with 0.9401±0.0047 after Agreement augmentation. Spatially blocked four-class experiments nevertheless showed that 20 to 30 local training labels per class recovered approximately 96–98% of the corresponding full-data performance. The framework therefore supplies an operational criterion for using unlabeled observations: augmentation is adopted only where calibrated filtering yields adequate class coverage, and validation confirms a downstream effect; otherwise the supervised model is retained. For the strict Sidi Moussa–Oualidia reduced-label experiment, the reported development budgets count every site-specific label used for fitting, early stopping, and calibration. The Merja Zerga blocked experiments separately quantify training-label sensitivity while retaining their blocked validation resources.