DOI: 10.3390/ijgi15100447 ISSN: 2220-9964

Separating Runtime Acceptance from Offline Contract Alignment in LLM-Orchestrated Flood GIS Workflows: A Dual-Case Evaluation

Hongyun Zhang, Jin Liu, Fang Wang, Yahong Zhao, Ye Xuan

Flood assessments require reproducible links between spatial products and the measurements ultimately reported to analysts, yet a large language model (LLM) agent may complete a geographic information system (GIS) workflow without generating the required product or traceable claim. We present GeoFloodAgent, an LLM-orchestrated flood-GIS workflow that combines typed planning and deterministic GIS tools with a separate evaluation of runtime disposition and post-run alignment with task-specific product, lineage, and action contracts. We evaluated 80 Chinese-language tasks from the Zhengzhou (2021) and Forlì (2023) flood cases under seven configurations and three repetitions, yielding a registry-remediated composite of 1680 assignments. In the Full configuration, 212 assignments (88.3%) were accepted and contract-aligned, whereas 16 accepted assignments did not align with their registered contracts; all 16 concerned clarification or refusal actions. Three runs were rejected despite retained contract-aligned evidence. The cluster-weighted Full-Tool Agent differences were −1.15 percentage points for accepted contract mismatch and +1.07 percentage points for contract-aligned acceptance, with empirical intervals spanning benefit and harm. Under frozen Zhengzhou inputs, deterministic outputs were sensitive to the minimum SAR object-size setting: relative to 80 pixels, candidate flood extent changed by +105.6% at 40 pixels and −30.9% at 120 pixels, whereas a 20 m simplification tolerance reproducibly failed during geometric union. These configuration-specific results do not validate hydrological accuracy or identify universal optimal defaults. These findings show that completion rate alone is inadequate for evaluating LLM-orchestrated flood GIS: runtime disposition and evidence alignment should be reported jointly. The study evaluates a finite, study-authored benchmark and does not establish hydrological accuracy, operational decision validity, or open-world error detection.