Integrating Sediment Geochemistry with Explainable Machine Learning for Provenance Discrimination in Wular Lake, Kashmir Himalaya, India
Mukhtar Hasan Ahmad, Shaik A. Rashid, Mohammad Khalid, Javid A. Ganai, Shamshad Ahmad, Amir Khan, AbuzarThis study integrates conventional sediment geochemistry with explainable machine learning to investigate the provenance of surface sediments from Wular Lake, Kashmir Valley, NW Himalaya. Twenty-two samples were analysed for 49 geochemical variables (10 major oxides, 25 trace elements and 14 rare earth elements), complemented by XRD mineralogy, which reveals an assemblage dominated by quartz, muscovite/illite, chlorite and feldspar. The Chemical Index of Alteration (CIA = 68.5–75.1, mean 72.1), corroborated by CIW, PIA and the A–CN–K trend, indicates moderate weathering under a cold temperate climate, and the Index of Compositional Variability (ICV > 1), together with uniformly low Zr/Sc ratios (3.4–6.0), which preclude significant zircon addition through recycling, records compositionally immature, first-cycle detrital input. Conventional discrimination ratios and the Herron system classify the sediments as geochemically equivalent to shale, and elevated Fe2O3/K2O (2.6–4.0), Al2O3/TiO2 (12.6–16.0), Cr/Th and Co/Th ratios record a substantial mafic imprint. Chondrite-normalised REE patterns show pronounced LREE enrichment ((La/Yb)N = 8.0–19.4), moderate negative Eu anomalies (Eu/Eu* = 0.56–0.73) and negligible Ce anomalies (Ce/Ce* = 0.98–1.03), with Eu/Eu* discriminating felsic crystalline from mafic volcanic contributions. A three-stage pipeline (principal component analysis (PCA) → random forest → SHapley Additive exPlanations (SHAP)) achieved a median leave-one-out cross-validation (LOO-CV) accuracy of 95.5% (n = 22; Wilson 95% CI 78%–99%), and unsupervised k-means clustering reproduced the same three geochemically distinct provenance end-members: siliceous-mature, detrital-mafic and carbonate-bearing, without reference to the assigned labels (Adjusted Rand Index = 1.0). Because the training labels derive from the same geochemical dataset, the classification quantifies the internal consistency of the provenance model, but does not provide independent validation. SHAP analysis reveals that trace elements (Cr, Co, Sc, Ni, Zn) carry greater discriminating power than do conventional major-oxide ratios, demonstrating that explainable machine learning robustly supplements and extends traditional provenance approaches.