DOI: 10.3390/math14152850 ISSN: 2227-7390

Stability-Regularized Residual Neural ODEs: From Rollout-Error Contraction Diagnostics to a Train-Time Method for Robust Long-Horizon Forecasting

Qin Li, Min Wan

Residual neural ordinary differential equations (NODEs) of the form f^=f+hθ can attain small one-step prediction error yet diverge under autonomous long-horizon rollout. A recent diagnostic attributes this to the one-sided Lipschitz (OSL) constant—the supremum over visited states of the logarithmic norm μ2(Jf^)=λmax(12(Jf^+Jf^⊤))—which, when positive, signals local expansion and amplifies persistent approximation error. In this work we convert this post hoc diagnostic into a train-time method by augmenting the one-step objective with a contraction penalty λEx[(μ2(Jf^(x))−c)+], and we study when this improves robust forecasting across stable, expansive, marginal, and chaotic regimes under realistic sensor-corruption noise. We prove that, at the penalty’s minimizer, the empirical OSL constant is controlled on the training set. A sample-to-domain covering condition then yields, via a Gronwall-type comparison, a conditional uniform-in-time rollout-error bound, with a time-averaged variant that justifies penalizing the mean rather than the maximum log-norm. Empirically, on a six-system, four-noise benchmark the penalty reliably drives the OSL constant down by one-to-two orders of magnitude, but whether this helps long-horizon accuracy is strongly regime-dependent and λ-sensitive: a common default (λ=0.1) over-damps and degrades rollout, whereas a calibrated λ≈0.01 helps only for measurably expansive baselines. Under a seed-decoupled re-evaluation with 15 seeds, the benefit is robust on a near-unstable, rotation-dominated oscillator—a 2.3× lower 100-step rollout error (p=0.018) at matched one-step error—but the apparent 5-seed improvement on a six-dimensional chemical reaction network does not replicate: with model and dataset seeds decoupled it reverses to a significant degradation, identifying the original effect as a seed-coupling artifact. The method thus yields a single robust positive result, and degrades already-contractive, conservative, and chaotic systems; for chaotic systems this is unavoidable, because enforced contraction suppresses the positive Lyapunov exponents that define the attractor. Finally, while a contraction-aware spectral penalty matches the log-norm penalty, standard ∥J∥2≤1 spectral normalization fails on rotation-dominated dynamics (8× worse rollout, p=0.0005, 15 seeds), confirming that the rotation-invariance of μ2 is the operative property.

More from our Archive