A Hybrid Machine Learning Framework for Multi-Pollutant Air Quality Assessment in Urban Environments

0Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

Urban air quality assessment is central to environmental sustainability and public health management. This study presents a structured comparative evaluation of Random Forest (RF), Support Vector Machine (SVM), LSTM, and Bi-LSTM models for pollutant-driven air quality classification under the Indian National Air Quality Index (NAQI) framework defined by CPCB guidelines. To provide a fair comparison, multi-pollutant data of Indian urban monitoring stations were preprocessed, and the class-balancing protocol and validation protocol were combined. RF had highest total accuracy (0.9971) in the held-out set, with Bi-LSTM (0.9615), LSTM (0.9495), and SVM (0.9442) coming next. Although ensemble methods proved to be very separable in line with the threshold-based NAQI structure, Bi-LSTM was more stable when it came to boundary-sensitive switches among the adjacent severity classes. Calibration analysis (multiclass Brier score: 0.08) showed consistent probabilistic behavior and interpretation, and using SHAP showed physically significant pollutant driving factors. The results explain the appropriateness of comparative models in organized AQI classification and present a reproducible assessment framework for the NAQI framework.

Cite

CITATION STYLE

APA

Mustafa, M., Akhtar, M., Ahmad, A., Javaid, F., Haldar, B., & Nisar, B. (2026). A Hybrid Machine Learning Framework for Multi-Pollutant Air Quality Assessment in Urban Environments. Sustainability (Switzerland), 18(4). https://doi.org/10.3390/su18042148

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free