DOI: 10.3390/electronics15194496 ISSN: 2079-9292

Forecasting of PM2.5/PM10 Using Machine Learning: A Benchmarking Study Based on Open Air Quality IoT Datasets

Ioannis Psomadakis, Christos Christakis, Christina L. Metallidou, Theodor Panagiotakopoulos, Angeliki I. Katsafadou, Yiannis Kiouvrekis

Particulate matter (PM2.5 and PM10) forecasting is increasingly framed as an applied machine-learning problem operating on real-world environmental sensor infrastructure, yet most benchmarking studies evaluate models on a single, curated station rather than on the heterogeneous, imperfect data that operational networks actually produce, and on a single chronological test split whose representativeness is rarely questioned. This study benchmarks four supervised model families, Random Forest (RF), Support Vector Regression with an RBF kernel (SVM), a feed-forward Neural Network (NN), and a Long Short-Term Memory network (LSTM), against a naïve persistence baseline and a conventional autoregressive baseline (SARIMA), across five of six monitoring stations of the Greek National Air Pollution Monitoring Network (EDPAR), a sixth, short-record rural station retained for exploratory analysis only, selected to span contrasting emission regimes and record lengths of 4 to 24 years. Using a univariate, past-only feature set (autoregressive lags, trailing rolling statistics, and cyclical calendar encodings), we forecast both next-day concentration and next-week maximum concentration for PM2.5 and PM10 independently at each station, under a two-stage evaluation design: a 40-window sliding model-selection stage, which selects each model family’s hyperparameters by mean R2 across many chronological cutoffs, and a single held-out 20% model-assessment stage on data never used for selection. SVM is the most broadly reliable model family under both stages, winning 11 of 20 station/pollutant/horizon combinations under model selection and 13 of 20 under model assessment, with its advantage most pronounced at the seven-day-maximum horizon; RF, NN and LSTM are each competitive at specific stations, but none is reliably dominant. The two evaluation stages agree on the winning model in only 10 of 20 combinations, illustrating that model rankings from a single chronological split can depend materially on the specific evaluation window chosen. SARIMA underperforms the best machine-learning model in all 20 of 20 combinations and underperforms naïve persistence itself at several stations, indicating that the machine-learning models’ advantage reflects genuine predictive skill rather than merely the seasonal and autoregressive information already available to any conventional time-series method. Achievable R2 is generally, though not universally, lower for PM10 than for PM2.5, reflecting PM10’s larger coarse-mode, episodically-driven component; station-specific exceptions to this pattern are attributable to long-term trend and test-set variance-compression effects rather than to any intrinsic reversal of pollutant forecastability. These results indicate that model selection for PM forecasting should be validated across multiple evaluation windows rather than a single split, that conventional statistical baselines remain a necessary comparison point, and that reported gains from more complex architectures should be interpreted cautiously in light of test-set non-stationarity and the sensitivity of model rankings to the choice of evaluation window.