CAML
‐
RNA
: Commutative Algebra Machine Learning for
RNA
–Ligand Binding Affinity Prediction
Stalin Arulsamy, Yashwanth Krishna, Rajesh Kumar, Vanktesh Kumar ABSTRACT
We introduce commutative algebra machine learning for RNA (CAML‐RNA), a framework that combines bipartite persistent homology (PH) and persistent Stanley–Reisner theory (PSRT) for predicting RNA–small‐molecule binding affinities. RNA–ligand interactions are represented as bipartite atom‐pair point clouds, enabling element‐specific (ES, 36 pairs, 4 RNA × 9 ligand elements) and category‐specific (CS, 36 pairs, 4 RNA structural categories × 9 ligand elements) Vietoris–Rips filtrations that extract β 0 and β 1 Betti curves (PH features, 3888‐dim combined), supplemented by persistent f ‐vectors and h ‐vectors from a true bipartite distance filtration (PSRT features, 3744‐dim, following Suwayyid and Wei). Applied to a curated benchmark of 143 RNA–ligand complexes spanning seven structurally distinct subtypes, CAML‐RNA achieves Pearson's R = 0.7283 (RMSE = 1.08 pK d units) under leave‐one‐out cross‐validation and R = 0.6255 under nested 10‐fold cross‐validation (), outperforming AffiGrapher ( R = 0.498), RLaffinity ( R = 0.559), and RLASIF ( R = 0.666, all LOO‐CV). A subtype‐aware feature selection strategy achieves R = 0.940 for aptamers and R = 0.771 for the riboswitch family.