DOI: 10.3390/su18199743 ISSN: 2071-1050

Building-Scale GeoAI for Short-Horizon Urban Inundation Prediction Under Sparse Labels and Distribution Shift: A Multi-Event Evaluation of Predictive Capability Boundaries in Taipei

Ming-Chih Jason Wang, Chien-Min Chen

Urban inundation forecasting in high-density cities is complicated by sparse incident labels, nonstationary rainfall, and changing sensing systems. Using Taipei as a case study, we develop a building-scale GeoAI decision-support prototype that integrates 30,421 Level of Detail 1 (LOD1) building-centroid nodes, 41,636 graph edges, built-environment attributes, and eight extreme-rainfall events from 2015 to 2026. We evaluate a static inundation-susceptibility baseline, 60 and 120 min event-level forecasts, tree-ensemble models, a Temporal GNN, and ConvLSTM and separately test temporal and spatial transfer under compound distribution shift. The analysis treats the data as having positive–unlabeled characteristics and distinguishes event-level, pooled, temporal-transfer, spatial-transfer, and scale-sensitivity evaluation regimes. All stochastic models were retrained under five fixed random seeds. Across the six primary events (the simulation-dominant E2016_Megi event excluded as a sensitivity case), the five-seed macro-averaged AUCs were 0.651 ± 0.035 at 60 min (95% event-cluster bootstrap CI 0.533–0.789) and 0.626 ± 0.016 at 120 min (0.445–0.805); including E2016_Megi raises the seven-event means to 0.694 ± 0.030 and 0.672 ± 0.014. Under non-equivalent spatial supports, the five-seed benchmark yields random forest AUC 0.711 ± 0.005, XGBoost 0.695 ± 0.000, Temporal LSTM 0.622 ± 0.009, Temporal GNN 0.477 ± 0.011, and ConvLSTM 0.408 ± 0.044; the ranking is configuration-specific and does not establish architecture superiority, whereas forward temporal validation falls to 0.549 ± 0.062. Physics-based synthetic augmentation increased the mean AUC for four historical gauge-era events only modestly, from 0.651 ± 0.013 to 0.668 ± 0.009, and one event deteriorated. These results define the conditions under which building-scale GeoAI remains informative and where it fails, providing a more defensible basis for urban-resilience decision support, climate adaptation, and infrastructure continuity.