DOI: 10.1021/acs.jcim.6c02473 ISSN: 1549-9596

Benchmarking TAS2R Subtype-Selectivity Prediction under Scaffold-Aware Evaluation: A Controlled Analysis of Receptor Representations

Jeoungyun Kim, Ku Kang, Jin Yoo

Abstract

Bitter taste perception is mediated by 25 human G protein-coupled receptors (TAS2Rs) whose ligand selectivity profiles range from receptors recognizing hundreds of chemically diverse compounds to receptors with fewer than 10 known agonists. Predicting which specific TAS2R subtype a compound activates─rather than simply whether it tastes bitter─is essential for understanding subtype-specific physiology and for designing selective bitter-masking agents. Here, we build a TAS2R subtype-selectivity model, TAS2R-SelectNet, and use it as a controlled testbed to establish two things about how such models should be evaluated and how their receptors should be represented. First, we quantify scaffold-based data leakage in commonly used TAS2R evaluation setups and identify its mechanism. Random-split evaluation inflates area under the curve (AUC) by up to 0.21 for receptor subtypes of broad selectivity (TAS2R39: ΔAUC = +0.210; TAS2R14: +0.188). The inflation tracks chemical similarity directly: prediction AUC rises from about 0.64 for the most novel test compounds to a plateau near 0.95 once a near-duplicate of a training compound exists, and random splits enrich the test set with exactly such compounds. Reported accuracies obtained under random splits are therefore likely optimistic, and the effect can be diagnosed in any benchmark with a simple similarity analysis. Second, we show that the widely assumed benefit of restricting protein language model embeddings to binding-pocket residues does not hold for this task, using controls that isolate what the receptor encoder actually exploits. A single residue drawn at random from outside the pocket performs as well as the pocket itself (0.735 ± 0.010 versus 0.721 ± 0.010 for a single-pocket residue), and for TAS2R14 the cholesterol-binding site and the intracellular agonist site are interchangeable (0.738 versus 0.739). At matched input dimensionality, full-sequence ESM-2 pooling (0.759 ± 0.004) leads pocket pooling (0.741 ± 0.004), with one-hot encoding between them (0.744 ± 0.011); the same holds for receptors held out entirely from training. Sweeping the receptor-encoder input width from 32 to 5120 dimensions across seven representations shows why such comparisons are fragile: all 42 settings fall within 0.708–0.771 macro AUC and the ranking changes with width, while training AUC reaches 1.000 throughout, so at this data set size, the model memorizes the training associations irrespective of receptor encoding. Comparisons between receptor representations therefore require matched input dimensionality and are not established by a single setting. We further test the model on held-out literature agonists absent from BitterDB, where it recovers the correct receptor for data-rich subtypes and fails for the near-orphan TAS2R2, delimiting the class-imbalance regime, and we release predictions and a selectivity index for all 14,835 compound–receptor pairs. TAS2R-SelectNet and all code are freely available at https://github.com/tysonyj/TAS2R-SelectNet.