DOI: 10.1021/acs.jcim.6c01839 ISSN: 1549-9596

Methodological Considerations in Small-Sample Multitask QSPR Modeling for PAMPA Permeability

Zekai Yu

Abstract

Formanek et al. reported a multitask PAMPA dataset of 143 drugs and drug candidates and compared multiple molecular descriptors and regression models for predicting passive membrane permeability. This Letter further examines two conclusions drawn from those results, with the aim of qualifying rather than overturning them. First, the external test set contains only 29 compounds, making small differences in test-set R2 difficult to interpret reliably. The test-set R2 and correlation metrics can also diverge because R2 is sensitive to scale and bias, whereas correlation primarily reflects ordering. In addition, membrane-specific experimental noise and substantial censoring at the detection limit further limit the robustness of rankings based on small differences in R2. Second, the Percepta descriptor set includes several predicted permeability and absorption endpoints, including LogPS, LogBB, Caco-2 permeability, and bioavailability. Its apparent advantage over learned molecular representations may therefore arise partly from these pretrained property predictions rather than from structural representation alone. We suggest compound-by-compound paired tests or bootstrap confidence intervals for small external test sets and representation comparisons in which these predicted endpoints are either removed from Percepta or equivalently provided to the learned representations.

More from our Archive