PFAS Prophet: Prioritising Per- and Polyfluoroalkyl Substances Using Mass and MS/MS Spectra
Kevin Thomas, Jake O'Brien, Pradeep Dewapriya, Saer Samanipour, Ian A. Wood, Mathieu François FeraudPer- and polyfluoroalkyl substances (PFAS) are a diverse group of over 14,700 synthetic chemicals known for their environmental persistence and association with adverse health effects. Non-target analysis (NTA) using high-resolution mass spectrometry (HRMS) is increasingly used for detecting and discovering new PFAS. However, manual analysis of HRMS data is time-consuming, and current automated systems depend on known reference standards and are restricted to Data Dependent Acquisition (DDA) methods. Here we propose a machine learning algorithm to prioritize potential PFAS-related compounds based on spectral data. The model, trained on distinct compound structures, can prioritize both known and unknown PFAS-related compounds without relying on prior statistical assumptions or reference standards, achieving an F1-macro score of 0.928, however only achieving a F1 score of 0.698 on PFAS related spectra. It leverages complex fragmentation patterns, KMD, and neutral losses to enhance detection accuracy.