Interpretable Classification of Missense Variants Using ESM2 Embeddings and UMAP
Ugo Lomoio, Tommaso Mazza, Pierangelo Veltri, Pietro Hiram GuzziAbstract
Missense variants can alter protein stability, conformation, activity, or expression, yet their clinical interpretation remains challenging, particularly for variants of uncertain significance. Protein language models provide sequence derived representations that capture evolutionary and biophysical constraints, but their high dimensionality limits direct interpretation. We present a framework combining residue level embeddings from the Evolutionary Scale Modeling protein language model with Uniform Manifold Approximation and Projection and supervised machine learning models for missense variant classification. The framework was evaluated across eight proteins associated with amyloidosis, monogenic diabetes, and cystic fibrosis. We compared three strategies: distance based classification in a two dimensional reduced embedding space, direct classification from the original embeddings, and classification from the dimensionality reduced embeddings. Performance was assessed using held out test data, discrimination metrics, calibration analysis, statistical comparison of receiver operating characteristic curves, and repeated runs with different random seeds. Predictive performance varied across diseases and evaluation metrics; supervised models based on the original embeddings provided the most consistent calibration, whereas the dimensionality reduction method offered an interpretable representation of protein specific variant landscapes. Overall, the proposed framework provides a scalable, sequence based approach for variant prioritization while explicitly separating visual interpretation from quantitative classification.