Beyond Plastic Labels: Polymer-Aware Substrate Encoding Improves Plastizyme Prediction and Reliability Models
Joseph Wasswa, Bharath RamsundarAbstract
Machine learning is increasingly being used to identify plastic-degrading enzymes (plastizymes), yet systematic evaluations of how different protein and polymer representations influence predictive performance, generalization, and computational cost remain lacking. Here, we systematically evaluated protein–polymer representation strategies using a curated data set of 720 enzyme–polymer interactions spanning seven polymer families. Handcrafted and learned protein and polymer representations, including transformer-based embeddings, graph representations, molecular descriptors, molecular fingerprints, and polymer language model embeddings, were evaluated individually and in multimodal combinations. Models were assessed using random, MMseqs2 sequence-clustered, and leave-one-polymer-family-out splits to quantify both in-distribution performance and out-of-distribution generalization. Under random splits, several multimodal models achieved excellent predictive performance (AUROC > 0.95), with transformer-based protein embeddings consistently outperforming conventional representations. Performance declined substantially under sequence-clustered and polymer-family-based evaluations, underscoring the difficulty of generalizing to remote enzymes and previously unseen polymer classes. Incorporating polymer representations consistently improved predictive performance, particularly for sequence- and descriptor-based protein models, while multimodal combinations provided the greatest benefit under the most challenging evaluation settings. Ablation and SHAP analyses revealed an increasing reliance on learned protein and polymer embeddings as model complexity increased. A quantitative cost-benefit analysis further showed that conventional polymer descriptors provided the best balance between predictive performance, model reliability, and computational efficiency, whereas increasingly complex learned polymer representations yielded only modest additional gains. Overall, this study provides a practical framework for evaluating protein and polymer representations and offers guidance for developing more robust, generalizable, and computationally efficient machine learning models for enzyme–polymer interaction prediction.