Soluble protein analog selection engine (SPASE)
: An automated
AI
‐powered server to improve protein engineering workflows
Sacha T. Larda, Alex Paré, Nicolas Doucet Abstract
The design of proteins with desired biophysical properties, such as high solubility and low aggregation propensity, is crucial for various biotechnological and biomedical applications. While deep learning‐based methods like ProteinMPNN have shown remarkable success in protein sequence design, their direct output may not always exhibit optimal solubility and aggregation properties. Here, we present Soluble Protein Analog Selection Engine (SPASE), a novel automated webserver that addresses this challenge by integrating ProteinMPNN with state‐of‐the‐art tools for protein solubility prediction (Protein‐Sol) and aggregation prediction (Aggrescan3D). SPASE automatically generates a diverse pool of protein variants using soluble ProteinMPNN, predicts the solubility of each analog, models their three‐dimensional structures with ESMFold, and scores these variants based on their predicted solubility, aggregation propensity, and folding confidence. Computational benchmarking indicates that SPASE enriches for protein analogs with higher predicted solubility and lower predicted aggregation propensity than the average output of soluble ProteinMPNN. We discuss the advantages and limitations of the workflow, including challenges associated with protein novelty, solubility prediction, and aggregation assessment. These considerations highlight the value of integrated platforms for prioritizing protein designs across multiple predicted biophysical properties. Together, these results position the SPASE server as a practical and accessible computational platform for prioritizing protein engineering candidates for downstream experimental evaluation. SPASE is publicly available at