Selective Prediction and the Persistence Illusion: A Diagnostic Decomposition of VIX Regime Classification
Akshat Gupta, Jianguo LiuThis paper develops a six-component diagnostic protocol for evaluating confidence-based selective prediction on autocorrelated financial labels, demonstrating it on VIX regime classification across 5000 trading days (July 2006–May 2026). The protocol combines coverage-matched baseline comparison, joint-coverage decomposition, regime-transition auditing, risk–coverage analysis, feature attribution, and abstention confound testing. Applied to Random Forest, Histogram Gradient Boosting, and XGBoost classifiers with 36 features, the protocol indicates that headline selective accuracy of 90–94% largely reflects label persistence: on jointly covered days, the Random Forest and a persistence rule calibrated on training data make identical predictions at three of five horizons (McNemar b=c=0) and disagree on at most 3 of 708 days elsewhere. All three classifiers and a forecast-then-threshold HAR model achieve 0% covered accuracy on calm-to-high transitions—across 16 distinct transition episodes at the shortest horizon—and an LSTM performs below chance on the same days. The pattern is consistent with an informational constraint of the daily feature set rather than a model deficiency, and it replicates on S&P 500 trend regimes. The protocol is modular and directly applicable to other selective prediction systems on serially correlated outcomes.