An Applied Assessment of Multi-Source Data Fusion by Machine Learning for PM2.5 Daily Concentration Prediction
Suhrudh Chivukula, Adrian J. Cortes Santos, Ruben Delgado, Dimuthu K. Arachchige, Jordan A. Caraballo-Vega, Mariel D. FribergAccurately predicting fine particulate matter (PM2.5) concentrations in regions with sparse monitoring networks remains a critical challenge for air quality management and public health. This study evaluates a machine learning (ML) data fusion approach that integrates daily federal regulatory observations, daily low-cost community sensor measurements, and monthly satellite-derived aerosol products (functioning as a regional background field) to improve PM2.5 prediction across under-monitored environments. Using a Long Short-Term Memory (LSTM) neural network architecture, the analysis examines how combining heterogeneous data sources influences predictions. Results show that pooled multi-source training was associated with higher holdout skill relative to some single-source configurations under this parsimonious baseline, though associations are city- and configuration-dependent and cannot be attributed solely to fusion because evaluation populations are not common. Comparisons against tree-based baselines (Random Forest, Gradient Boosting, XGBoost) indicate that overall predictive skill, not just the LSTM’s, is constrained by data availability, suggesting that data composition, rather than model choice, is the primary driver of the observed performance patterns. These findings highlight both the potential and the practical constraints of multi-source ML approaches for air quality prediction and exposure assessment, with implications for model design, monitoring strategy, and environmental equity. This study is intentionally scoped as an applied evaluation of data fusion performance rather than a comprehensive assessment of algorithmic optimality or operational forecasting readiness. The analysis focuses on daily PM2.5 prediction across a selected set of U.S. cities and does not address sub-daily variability, real-time deployment constraints, or event-specific model optimization. Model performance is therefore interpreted in the context of data availability, consistency, and representativeness, rather than as an upper bound on achievable predictive skill.