DOI: 10.3390/rs18183228 ISSN: 2072-4292

Region and Boundary Accuracy in Zero-Shot Shoreline Extraction: A Benchmark of Segmentation-Native Vision-Language Models and Prompted SAM Pipelines

Ge Shi, Haoran Tang, Boyuan Lu, Lin Sun, Taiyu Ye, Wei Wang

Vision-language models (VLMs) can segment object categories absent from their training data, offering a route to shoreline extraction that requires no site-specific annotation. Their accuracy in localizing the land–water boundary, and its dependence on prompt formulation, have not been established. This study evaluates four models representing two distinct pipeline families on 3104 aerial images from the Coastal Aerial Imagery Dataset: direct mask prediction models (CLIPSeg, Florence-2) and VLM-coordinate-guided SAM pipelines (GPT-4o, Qwen-VL-Max). Three prompt conditions were compared under two region-level and two boundary-level metrics, stratified by shoreline complexity, shoreline density, coastal landscape, and wave state. Expanding the prompt with detailed rules for discriminating water from land reduced accuracy relative to a minimal instruction in fifteen of sixteen model–metric combinations, with Florence-2’s mean boundary error rising from 107.08 to 120.98 pixels. A flexible anchoring strategy, adding a second positive point and making the negative point optional, raised region accuracy for all four models but improved boundary accuracy only for the segmentation-native pair. Region-level and boundary-level metrics were weakly associated (IoU–AED Spearman correlation between −0.20 and −0.47), so high overlap scores did not imply accurate boundary placement. Complete failures occurred in 11.2% to 27.4% of images, concentrated on sheltered shoreline settings with high shoreline density rather than wave-disturbed open coast. Although the median boundary error (21.86 pixels for CLIPSeg) suggests that successful zero-shot predictions approach the 2.71–29.55 pixel mean error range of supervised networks, overall reliability remains compromised by long-tail failures, with region accuracy trailing supervised baselines by 3 to 18 IoU points. Zero-shot extraction thus remains less accurate than supervised segmentation, but offers a label-free alternative whose reliability may be improved through architecture-aware prompt design and environment-aware quality control.