Empirical Post-Generation Correctness Assessment for Natural-Language Access to Urban PostGIS Databases
Jiqiu Deng, Chaolin Zhang, Hui Zhang, Longbo Li, Liji Sun, Xiao Ma, Zhiyong GuoNatural-language interfaces make urban spatial databases accessible, but generated structured query language (SQL) can execute successfully while returning a semantically incorrect answer. We evaluate a post-generation correctness-ranking layer on 18,900 candidates from Nanjing, Wuhan, and Shenzhen. The original city split is a controlled, highly template-aligned benchmark rather than a structure-disjoint deployment test. Under strict source-side separation of fitting, isotonic calibration, threshold selection, and target evaluation, Spatial-safe improves executed non-empty area under the receiver operating characteristic curve (AUROC) in all nine transfer directions by 0.062 on average. Grouped-threshold area under the risk–coverage curve (AURC) is favorable in 8/9 directions, and calibration is mixed. Source-selected thresholds raise mean conditional target coverage within the executed non-empty operating subset from 0.457 to 0.669, while mean empirical risk rises from 0.110 to 0.136. Unseen-template and SQL-skeleton-disjoint effects are smaller and mixed. To determine whether the existing benchmark could support the adequately powered structure-disjoint comparison, we applied a predeclared power/data-sufficiency feasibility gate. With only 77 independent SQL-skeleton groups, 42/60 required cells fail the gate; therefore, the present benchmark cannot support an adequately powered structure-disjoint performance comparison without additional skeleton-diverse questions. We do not substitute an underpowered estimate for that missing evidence. The evidence therefore supports empirical correctness-ranking improvement in this controlled benchmark, not universal structure-disjoint generalization, formal SQL verification, benchmark-wide human validation of all correctness labels, or target-domain risk guarantees.