PC-Audit: A Decision-Reliability Framework for Auditing Validation-Based Machine-Learning Model Selection in Weak-Signal Financial Time Series
Songrun Li, Weiran Zhang, Yichen Wang, Qi LeiValidation-based machine-learning selection can name a numerical winner in temporally dependent weak-signal data without establishing that the decision will remain reliable in a later period. We present PC-Audit (Prequential-Calibration Audit) as a decision-reliability architecture—not a new forecasting model or accuracy-enhancement method—that records the evidence surrounding a validation-selected candidate. It combines target-date partitioning, prequential non-negative calibration, Model Confidence Sets (MCS), transfer and refit diagnostics, dependence sensitivity, and forward-only governance records. In 45 walk-forward cases for nine Chinese index ETFs (2020–2024), signed-return validation ranking transfers weakly (mean rank correlation 0.090), PC-Select achieves the lowest test MAE in 15 of 45 cases, and the prequential MCS retains 4.80/5 candidates on average. L1 calibration sensitivity, alternative validation horizons, and dependence sensitivity do not support a strong signed-return selection claim. Base-model refitting adds ranking instability to already weak validation-to-future transfer. A frozen controlled simulation shows that the protocol is not mechanically abstention-only under clearly separable stationary oracle risk, while an unobserved future reversal remains a prospective limitation. Daily information-aligned econometric comparators and QLIKE analyses show that audit evidence is target- and loss-dependent: squared-return results can improve under QLIKE-aligned selection, whereas absolute-return and range-proxy results do not improve uniformly. Fallback uncertainty intervals cross zero. PC-Audit therefore supports transparent qualification of validation-selected candidates and completed-period audit/next-cycle governance, rather than deployment authorization or claimed predictive improvement.