Machine Learning for Colloidal Stability and Aggregation Risk in Biopharmaceutical Formulations: Evidence, Limits, and Practical Use
Carlos Victor Montefusco-PereiraMachine learning is increasingly used to relate molecular descriptors, formulation variables, and biophysical measurements to aggregation, viscosity, solubility, and shelf-life outcomes. The evidence is promising but uneven. Most published datasets contain tens to a few hundred antibodies, use different assays and endpoint definitions, and rely mainly on internal validation. Direct evidence for bispecific antibodies, antibody–drug conjugates, mRNA–lipid nanoparticles, and viral vectors remains limited. This structured critical review evaluates what current models can support, how data and validation choices shape reported performance, and where claims exceed the available evidence. We searched PubMed through 30 June 2026 using predefined queries for machine learning, biopharmaceutical formulation, colloidal stability, advanced modalities, and shelf-life modelling. Studies were assessed by molecular diversity, formulation coverage, endpoint quality, split strategy, external validation, and decision relevance. The strongest current use cases are early antibody developability screening, high-concentration viscosity classification, formulation ranking within a defined experimental domain, and image-based particle classification. Long-term shelf-life prediction may benefit from hybrid kinetic and machine learning models, but real-time confirmation remains necessary. Progress will depend less on larger algorithms than on better labels, molecule-level validation, shared reference datasets, and clear uncertainty reporting.