DOI: 10.1111/jpg.70099 ISSN: 0141-6421

Machine Learning Prediction of Cambrian Source Rock Quality With Calibrated Uncertainty: A Case Study From the Qiongzhusi Formation, Upper Yangtze Platform, South China

Basel Alrawi, Xiaoping Mao, Wenhui Huang

ABSTRACT

Predicting source rock quality from geochemical proxies is fundamental to shale gas exploration, yet conventional approaches provide point estimates without quantifying prediction confidence. We develop and validate ensemble machine learning models for predicting total organic carbon (TOC) and source rock quality across the Lower Cambrian Qiongzhusi Formation, using 1264 core samples from 12 wells spanning 5 depositional settings, and introduce conformal prediction to deliver calibrated uncertainty intervals. Random Forest regression performance is reported across three operationally distinct configurations: trace‐element geochemistry alone predicts TOC at R 2  = 0.77 (MAE = 0.95 wt%), incorporating depositional setting raises this to R 2  = 0.88 (MAE = 0.71 wt%), and the four‐feature reduced model (Mo EF, V/(V + Ni), S, depositional setting) recovers R 2  = 0.87 (MAE = 0.74 wt%) from operationally minimal inputs. The reduced model is recommended as the operational tool for routine exploration screening. The full model substantially outperforms multiple linear regression ( R 2  = 0.78). Vanadium concentration is the strongest single predictor (impurity importance 50.4% and permutation importance 36.4%), with the contrast against the V/(V + Ni) ratio (0.6%) explained by the Mo reservoir effect in restricted‐basin settings. Conformal prediction provides calibrated uncertainty intervals: the 90% prediction interval is ±1.87 wt% TOC with empirical coverage of 90.0%, validated under both out‐of‐fold and strict two‐stage split procedures. Uncertainty varies systematically with depositional setting—from ±0.50 wt% on the platform to ±2.37 wt% in basin center settings. Binary classification at the operationally relevant TOC ≥ 2 wt% threshold for shale‐gas prospectivity achieves 98.5% accuracy. Leave‐one‐well‐out cross‐validation confirms transferability within sampled environments ( R 2 up to 0.85), but leave‐one‐setting‐out reveals systematic extrapolation failure to unsampled depositional environments (basin center R 2  = 0.36; platform R 2  = −2.6). These results indicate that calibrated uncertainty quantification, not point prediction, is the appropriate paradigm for ML‐based source rock assessment and that prediction models must be trained on data spanning the full range of target depositional environments to avoid systematic extrapolation failure.

More from our Archive