Reliable Apparent Permeability Prediction Through Benchmarking, Stability, Uncertainty Quantification, and Explainable Machine Learning
Rajdeep Mondal, Rajith K. R. Rajoli, Andrew Owen, Soumitra SamantaABSTRACT
Apparent permeability () of a drug molecule serves as a principal indicator of drug absorption for pharmacokinetic modeling, and accurate prediction of values using machine learning would avoid time‐consuming and expensive In Vitro experiments. In this study, four different feature representations were compared for predicting values, with each representation serving as input to three subsequent regression models. The best predictive model achieved an value of 0.68 and RMSE of 0.42 on the test data. Furthermore, the SHAP values corresponding to individual molecular descriptors were analyzed to interpret the influence of these descriptors on the model prediction. Some key properties like octanol–water partition coefficient and number of basic atoms were found to have a strong influence on the value. Additionally, some specific molecular substructures were identified that contribute either positively or negatively to the model output, thereby inferring effects on drug permeability. Finally, the uncertainty in the predictions was quantified through a distribution‐free and model‐agnostic method, jackknife+ , used for the estimation of confidence intervals. At the optimal experimental setting, the confidence intervals derived using the overall best performing model covered 96.12% and 86.73% of the original values with more than 95% confidence for test and independent data, respectively. With the same confidence, the confidence intervals contained 99.87% and 100% of the predicted values for the test and independent data, respectively, underscoring the reliability of the model. Overall, the presented ML provides the most robust informed prediction among compared models presented to date.