A Symmetric Prediction–Explanation Framework for Explainable Risk Warning: Integrating Large Language Models as an Interpretive Layer with Ridge Regression
Ke He, Xuefeng Xia, Changfeng WangFrom the perspective of functional symmetry design, this study constructed a functionally complementary framework between “prediction” and “explanation” in the field of risk early warning. To address the semantic gap between prediction results and decision-making needs, this study proposes a symmetric prediction–explanation collaborative framework. In this framework, the ridge regression serves as the prediction core, and a large language model (LLM) enhanced by retrieval-augmented generation (RAG) technology is introduced as a post hoc explanation layer. By dynamically retrieving the domain knowledge base, the framework generates accurate and credible semantic reports for risk warning. Ridge regression was compared with five baseline models, namely ordinary least squares (OLS) regression, least absolute shrinkage and selection operator (LASSO), elastic net, random forest, and linear support vector regression (SVR). The nested leave-one-out cross-validation (LOOCV) results showed MAE = 0.067, RMSE = 0.082, R2 = 0.836, MAPE = 12.9%, SMAPE = 12.4%, and a Spearman correlation coefficient of 0.892. Ridge regression significantly outperformed OLS and LASSO (p<0.05), but the differences with linear SVR, random forest, and other models did not reach statistical significance. Under the condition of a small sample size, the core value of ridge regression lies in its linear interpretability, rather than a marginal advantage in predictive accuracy. The RAG-enhanced LLM successfully transformed the safety assessment scores output by the model into a professional report containing a risk situation overview, key risk factor analysis, and preliminary action recommendations. The Recall-Oriented Understudy for Gisting Evaluation (ROUGE) automatic evaluation showed that under the RAG condition, the F1 scores for ROUGE-1, ROUGE-2, and ROUGE-L are 0.695, 0.442, and 0.523, respectively, representing improvements of 34.4%, 68.1%, and 68.2% over the condition without RAG. Expert evaluation further confirmed that the overall score of the report under the RAG condition is 0.84, compared to only 0.72 without RAG, verifying the critical role of RAG in suppressing hallucinations and improving factual accuracy. Through the functional synergy of ridge regression and LLMs, this study achieved information consistency between “prediction” and “explanation” in risk early warning, providing a feasible pathway for deploying trustworthy AI decision support systems in high-risk domains.