DOI: 10.3390/w18151902 ISSN: 2073-4441

Integrating Hydrochemistry and Explainable Machine Learning for Groundwater Quality Assessment in the Bismil Plain, Türkiye

Sevgi Özgür Geter, Süreyya Betül Rufaioğlu, Ali Volkan Bilgili, Güzel Yılmaz

This study evaluates groundwater quality in the Bismil Plain (Diyarbakır, Southeast Türkiye) using a total of 208 samples collected from 26 wells during eight seasonal sampling periods conducted between 2022 and 2024. In each sample, pH, electrical conductivity (EC), and the major ions Ca2+, Mg2+, Na+, K+, Cl−, SO42−, HCO3− and NO3− were analyzed, and a WHO-based Water Quality Index (WQI) was calculated for every observation. The study combines classical hydrochemical interpretation methods, including descriptive statistics, hierarchical correlation analysis, variance inflation factor, and Piper and Gibbs diagrams, with an explainable machine learning framework integrating SHAP-based feature selection into Random Forest, XGBoost, support vector regression, and stacking ensemble models. In addition, spatial residuals were evaluated using Moran’s I and ordinary kriging, anomalies were identified using Isolation Forest and Local Outlier Factor algorithms, and predictive uncertainty was quantified through bootstrap resampling. WQI values ranged from 79.37 to 125.48 (mean: 99.67), with all samples classified only within the “Good” (49.5%) and “Poor” (50.5%) quality categories, indicating that the aquifer is close to a critical water-quality threshold. Spatially, the highest (poorest-quality) WQI values form a coherent zone in the south-western and central parts of the plain, whereas the central-eastern wells return the lowest values; the same pattern is reproduced by all four models. XGBoost and the stacking ensemble models showed comparable predictive performance (R2 = 0.911 and 0.910; RMSE = 3.29 and 3.27, respectively), while SHAP analysis identified EC as the dominant controlling factor, followed by NO3−, SO42−, Ca2+, Mg2+ and Cl− (mean |SHAP| = 6.86, 1.57, 1.10, 1.09, 0.85 and 0.72 WQI units, respectively). Moran’s I computed on the residual fields was −0.067 (p = 0.275) for XGBoost and −0.068 (p = 0.273) for the stacking ensemble, so ordinary kriging of these residuals produced an essentially null correction, whereas the SVR residuals remained spatially autocorrelated (I = 0.242; p = 0.001) and were meaningfully corrected by the geostatistical step. The originality of the study lies in integrating explainable machine learning, geostatistical residual analysis, anomaly detection, and bootstrap-based uncertainty assessment within a unified framework for a multi-season groundwater dataset, while also evaluating the effectiveness of spatial correction using a Moran’s I-based approach.

More from our Archive