DOI: 10.3390/commodities5030016 ISSN: 2813-2432

Predicting Commodity ETF Returns with Deep Learning: Overnight Versus Daytime Predictability Across Forecast Horizons

Triparna Kundu, Sarthak Pattnaik, Eugene Pinsky

Commodity prices are notoriously hard to forecast, and whether the returns of commodity exchange-traded funds (ETFs) can be predicted remains an open question. We compare three deep learning models, Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), and Transformer, for forecasting the returns of six Deutsche Bank commodity ETFs covering agriculture (DBA), base metals (DBB), broad commodities (DBC), energy (DBE), oil (DBO), and precious metals (DBP). Using daily price data from January 2007 to December 2025, we predict daytime returns (open to close) and overnight returns (previous close to open) separately, over five horizons of 1, 5, 30, 60, and 180 trading days. Each model sees a 20-day window of price-based features, returns, rolling averages and volatilities, momentum, and recent lags, built from all six ETFs. All models are trained on a strict chronological split and judged by two simple, decision-oriented measures: how often they call the direction correctly, and the risk-adjusted return (annualized Sharpe ratio) of a stylized long–short strategy that ignores transaction costs. Formal significance tests with HAC corrections for overlapping targets, bootstrap confidence intervals, and comparisons with ARIMA, random forest, and simpler benchmarks corroborate strong predictability in overnight DBP and daytime DBB at medium horizons. Predictability turns out to be highly specific to the asset, the trading session, and the horizon. Overnight returns of the precious metals ETF (DBP) are by far the most predictable: the correct direction is called 71.6% of the time at 60 days and 76.7% at 180 days, with Sharpe ratios reaching about 15. Base metals (DBB) daytime returns are predictable at 30 days and oil (DBO) daytime returns at 180 days, whereas one-day-ahead forecasts and agricultural returns (DBA) stay essentially unpredictable. The Transformer has a slight edge at longer horizons and the GRU at shorter ones. Key directional accuracy and Sharpe ratio results are confirmed by Newey–West HAC significance tests and Diebold–Mariano forecast comparison tests with the Harvey–Leybourne–Newbold small-sample correction; HAC standard errors at the 180-day horizon exceed naïve OLS errors by a factor of approximately 7.4, and we explicitly flag results that do not survive this correction. A three-fold expanding walk-forward validation scheme corroborates the main findings, with DBP overnight and DBO daytime predictability persisting across all evaluation windows. Deep learning architectures statistically and economically outperform logistic regression, ridge regression, and momentum baselines on the most predictable configurations. An anomalous failure of all models on DBA daytime returns at the 180-day horizon is diagnosed as a regime-driven artefact associated with post-2021 commodity inflation, not a general feature of agricultural return dynamics. The broader lesson is that splitting returns into daytime and overnight components exposes predictable structure that conventional close-to-close returns hide.

More from our Archive