Abstract
Water quality in Peru is an increasing concern, particularly in the upper Huarmey watershed, which is affected by heavy metal contamination and untreated wastewater. This study proposes an automated classification approach using three supervised machine learning algorithms—K-Nearest Neighbors (KNN), Support Vector Machine (SVM), and Random Forest (RF)—to assess the water quality based on the Water Quality Index (WQI) of Peru. The experimental results show that KNN outperforms other methods, reaching an accuracy of 95.2%. The proposed system automates and improves the classification accuracy compared with manual methods based on Microsoft Excel. The methodology, performance metrics, dataset characteristics, and geographical context are detailed to ensure replicability. This algorithm assists decision-makers with environmental monitoring and public health protection.
Author supplied keywords
Cite
CITATION STYLE
Vega-Huerta, H., Pajuelo-Leon, J., De-la-Cruz-VdV, P., Calderón, D., Maquen-Niño, G. L. E., Rios-Castillo, M. E., … Benito-Pacheco, O. (2025). K-Nearest Neighbors Model to Optimize Data Classification According to the Water Quality Index of the Upper Basin of the City of Huarmey. Applied Sciences (Switzerland), 15(18). https://doi.org/10.3390/app151810202
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.