DOI: 10.1021/acs.analchem.6c01825 ISSN: 0003-2700

Foundation Models for Liquid Chromatography–High-Resolution Mass Spectrometry: A New Era beyond Labeled Datasets

Andrea Junior Carnoli, Federico Padilla-Gonzalez, Leonieke M. van den Bulk, Daan Korporaal, Martin Alewijn, Marco H. Blokland, Bas H.M. van der Velden

Abstract

Liquid chromatography coupled to high-resolution mass spectrometry (LC–HRMS) is a widely used analytical technique for characterizing the chemical composition of organic samples. Due to its high sensitivity and ability to detect thousands of chemical features in a single run, untargeted LC–HRMS experiments generate highly complex and data-rich datasets that typically require advanced computational methods, including machine learning, for meaningful interpretation. While traditional machine learning approaches have been applied to LC–HRMS data, their performance remains limited for complex tasks. Deep learning has demonstrated improved performance, but both machine and deep learning are often constrained by the complexity and scarcity of labeled LC–HRMS data. Foundation models present a promising new horizon for LC–HRMS data analysis, given their ability to learn transferable representations from large-scale unlabeled data and adapt efficiently to downstream tasks with limited labeled samples. Recent studies have shown that foundation models can outperform conventional machine learning approaches in chemical annotation and molecular property prediction. We envision that foundation models for LC–HRMS data will benefit from the expansion of curated sample repositories and spectral libraries, developing privacy-preserving training strategies, enabling simultaneous modeling of multiple LC–HRMS data types, and improving model explainability.

More from our Archive