Abstract
Several studies developed machine learning-based PM2.5prediction models; however, nationwide models addressing both mapping prediction and forecasting were limited. Further, although the prediction accuracy is different from PM2.5-related health risk estimation, previous studies solely examined the prediction accuracy. This study suggests a method to assess the statistical properties of PM2.5-health risk estimation, which also can be used as a model selection. We used three machine learning algorithms and an ensemble method to construct PM2.5mapping prediction (1 km2) and two-day forecasting models majorly using satellite-driven data in South Korea (2015–2022). We performed a simulation study to examine the statistical properties of short-term PM2.5risk estimation using prediction models. Our ensemble spatial prediction model showed better performance than single algorithms (0.956 test R2). The range of the R2values was 0.78–0.98 across the monitoring sites. The average % bias was from 1.403%–1.787% when our mapping models for PM2.5-mortality risk estimation, compared to the estimates from monitored PM2.5. The best R2of our forecasting models was 0.904. This study developed machine learning models for spatial PM2.5predictions and forecasting in Korea. This study also suggested a method to address risk estimation and model selection concurrently when multiple prediction models were used.
Author supplied keywords
Cite
CITATION STYLE
Ahn, S., Kim, A., Chung, Y., Kang, C., Kim, S., Kwon, D., … Lee, W. (2025). Nationwide Machine Learning-Ensemble PM2.5Mapping Prediction and Forecasting Models in South Korea with High Spatiotemporal Resolution and Health Risk Estimation-Based Evaluations. Environment and Health, 3(8), 878–887. https://doi.org/10.1021/envhealth.4c00201
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.