A Hybrid Transfer Learning Approach for Ambient PM2.5 Exposure Modeling in LMIC Contexts: A Case Study from Lima, Peru
Qiang Pu, Qingyang Zhu, Elvis Medina, Alan Llacza, Bryan Vu, Jianzhao Bi, Kenan Li, Laura Nicolaou, William Checkley, Stella Hartinger, Kyle Steenland, Yang LiuAbstract
Exposure to fine particulate matter (PM2.5) contributes to millions of premature deaths annually worldwide. While epidemiological studies increasingly rely on geospatial technologies, such as satellite remote sensing, low-cost sensors, and chemical transport models, to quantify air pollution exposure, sparse monitoring in many low- and middle-income countries often introduces sampling bias, potentially skewing exposure estimates. We developed a hybrid and transfer learning approach in Lima, Peru, that integrates multiple geospatial technologies and machine learning to mitigate sampling bias and improve PM2.5 predictions. Our approach incorporated ground PM2.5 observations from reference-grade and low-cost monitors, satellite-retrieved aerosol optical depth, simulations from Weather Research and Forecasting model coupled with Chemistry, and additional pollution-related variables. We accounted for the sampling bias using a geo-envrionmental similarity-based transfer learning, leveraging training data from data-rich regions (the United States) to inform the data-poor target region (Lima). We then built an instance-weighted random forest model to predict daily PM2.5 concentrations across Lima from 2010 to 2023. The model achieved a cross-validated R2 of 0.76 and reduced potential overestimation in mountainous regions (average PM2.5 reduced from 34.19 to 29.23 μg/m3). This approach provides a scalable solution to support robust epidemiological research and public health planning across data-scarce settings.