DOI: 10.3390/ijms27198784 ISSN: 1422-0067

Machine Learning-Based Classification of Pairwise Selectivity Among Human Carbonic Anhydrase I, II, IX, and XII Isoenzymes

Elif Gozde Hacisuleyman, Ahmet Kayraldiz, Erol Eroglu

Human carbonic anhydrase (hCA) isoenzymes participate in essential pH-regulatory processes, but conservation of their catalytic sites makes selective inhibition difficult. We developed a direct pairwise machine learning framework for selectivity among hCA I, II, IX, and XII using curated Ki data from ChEMBL. Compounds entered a binary task only when one isoenzyme was strongly inhibited (Ki ≤ 200 nM) and the paired isoenzyme was weakly inhibited (Ki > 800 nM). A provenance-conservative STRICT analysis and a BROAD sensitivity analysis were evaluated under train-only split selection, hyperparameter optimization, and a globally locked outer-test protocol. The highest observed held-out performance was obtained for CA II/CA XII (MCC = 0.748) and CA IX/CA XII (MCC = 0.718), whereas CA II/CA IX was intermediate (MCC = 0.459) and CA-I-containing tasks were generally more difficult. Because the highest point estimates arose from relatively small test cohorts, these findings should be regarded as promising and require independent external or prospective validation. All six STRICT models exceeded 100 randomized-label controls (empirical p = 1/101 ≈ 0.0099, the resolution floor for 100 permutations). Applicability domain and SHAP analyses indicated pair-specific signals compatible with known isoenzyme-dependent active-site differences. Overall, the framework prioritizes interpretable hCA selectivity hypotheses while explicitly defining current uncertainty and validation limits.