The Evaluation Gap: Why Standard Accuracy Metrics and Drift Detectors Understate the Deployment Risk of Crash-Severity Prediction Models
Thaar Alqahtani, Fawzan AlfawzanRoad authorities increasingly use machine-learning models to rank high-risk locations and to allocate scarce safety resources. These models are almost always judged by their predictive accuracy on a random hold-out sample. This study shows that practice hides an important risk. We evaluate crash-severity models the way they are used in practice. Each model is trained on past data and applied to later years. We measure decision quality, not only accuracy. The data include 45,288 police-reported crashes recorded across all thirteen administrative regions of Saudi Arabia between 2016 and 2021. Four results follow. First, decision quality decays much faster than accuracy. Over a three-year deployment gap, the area under the curve falls by less than two percent. Yet the overlap between the predicted and the true high-risk sites falls from 0.44 to 0.24, close to the value expected by chance. Second, the useful shelf life of a model is short. Site rankings approach random within about two to three years. Third, retraining recovers only a small part of the lost decision quality, at most about three percentage points of overlap. Fourth, the standard drift detector, the Population Stability Index on the score distribution, never crosses its alert threshold during this decay across eighteen design settings. A decision-aware monitor based on turnover in the model’s own high-risk set instead reveals churn of 0.20 to 0.60 from the same deployment-time data. Fifth, model selection also changes under decision-level evaluation. Ranking candidate models by accuracy and by decision quality disagree, with a rank correlation near zero, so the most accurate model is not the best decision-level performer. Standard accuracy-and-drift practice therefore misleads at both the selection and the monitoring stage. Crash-severity models should be evaluated and monitored at the level of the decisions they inform.