Prior-Informed Directed-Lag Graph Neural Residual Learning for Multi-Step Streamflow Forecasting
Liang Mu, Zhiguo Yu, Hongmin Zhang, Junxi ChenAccurate multi-step streamflow forecasting requires models to preserve antecedent streamflow memory and account for delayed dependence between gauges. This study proposes a prior-informed directed-lag graph neural residual model, termed PI-DLGNR. The framework decomposes prediction into a dominant linear memory component and a nonlinear residual correction regularized by soft routing priors. A multivariate VAR-Ridge backbone captures autoregressive persistence and cross-station streamflow memory. A directed-lag graph residual branch learns additional corrections using an inferred directed graph with trainable weights, learnable lag kernels, and station-wise temporal encoders. Three penalties regularize downstream ordering, hydrograph curvature, and delay structure without enforcing water balance. In the original benchmark on the GloFAS reanalysis series in the Yangtze River Basin, PI-DLGNR achieves NSE values of 0.998, 0.975, and 0.900 at Steps 1, 3, and 7, respectively. Relative to the isolated VAR-Ridge backbone, MAE decreases by 1.47%, 0.84%, and 0.72%, while RMSE decreases by 0.34%, 0.03%, and 0.01%. VAR-Ridge retains higher KGE at all three steps. In a separate three-seed comparison selected using chronological validation, the MAE-difference confidence intervals span zero at all three steps. The residual advantage is therefore not robust to the revised selection protocol. Tests with independent retraining in the Pearl and Yellow River mainstreams assess the robustness of the modeling strategy, not parameter transferability. Despite an overall Step 7 NSE of 0.900, PI-DLGNR has a Flood-NSE of −0.514, with negative Flood-NSE for every evaluated model. The evidence is stronger for general streamflow variation and selected low-flow conditions than for extremes one week ahead, although low-flow gains are not consistent across refits.