DOI: 10.1021/acs.analchem.6c02844 ISSN: 0003-2700

LC-HRMS Known- and Unknown-Feature Prioritization Based on Machine-Learning-Predicted Ready Biodegradability

Tobias Hulleman, Saer Samanipour, Viktoriia Turkina, Jinglong Li, Cassandra Rauert, Elvis D. Okoffo, Kevin V. Thomas, Jake W. O’Brien

Abstract

The vast and expanding chemical universe contains hundreds of thousands of substances, yet experimental data on their environmental persistence are available for only a small fraction (<3%). This gap limits our ability to identify persistent organic contaminants. Models have been developed to predict the persistence for already identified chemicals, but no models are able to prioritize unknown features detected in high-resolution mass spectrometry (HRMS). Here, we present two complementary machine-learning models that connect molecular structure and tandem mass-spectrometry fragmentation behavior to ready biodegradability. The first integrates molecular fingerprints with synthetic and retrosynthetic accessibility and molecular complexity metrics, achieving a cross-validated balanced accuracy of 88% for 4,840 curated chemicals. The second model extends predictions to unknowns by using cumulative neutral losses (CNLs) derived from HRMS/MS data and trained in a semisupervised manner guided by the structure-based model, reaching 82% balanced accuracy. Analysis of key CNLs revealed fragmentation patterns linked to biodegradability, such as hydrolytic- and halogen-related losses. Together, these models provide a data-driven approach to prioritize persistent known and unknown chemicals in complex environmental samples. When applied to HRMS/MS data from sewage-biodegradation reactors, the models captured trends consistent with observed degradation behavior, with readily biodegradable compounds decreasing in abundance and persistent compounds remaining stable over time.