Abstract
In machine learning (ML), class imbalance is a serious problem when datasets with a dominating majority class make it difficult to accurately classify minority samples. Traditional methods like the Synthetic Minority Oversampling Technique (SMOTE) may lead to producing the noise and overfitting. Consequently, model performance degrades. In order to address this problem, we develop an innovative technique Adaptive K-Nearest Neighbor-Based Oversampling (AKS), a novel approach that strategically generates synthetic minority samples using an adaptive neighborhood selection and non-linearity checks among the considered minority samples. AKS avoids to generate synthetic points in majority regions, which attempts to reduce over-fitting problem. Further, the proposed model is applied on ten well known imbalanced datasets to make them balanced. To test the utility of the balanced dataset, seven ML algorithms are applied on such datasets. To validate the usefulness of the proposed method, SMOTE, SMOTE-ENN and borderline-SMOTE are being utilized. Moreover, the models performance is validated through 5-fold cross-validation methodology. In addition, statistical analyses, including the Friedman test are conducted to rigorously evaluate and validate the robustness of the proposed method. The results shows that, the proposed technique performs at par or better than other data balancing techniques.
Author supplied keywords
Cite
CITATION STYLE
Nandini, A., Mishra, T. K., Pujahari, A., & Sahoo, K. S. (2026). AKS: Optimized Adaptive K-Nearest Neighbor-Based Oversampling for Imbalanced Datasets. IEEE Access, 14, 22699–22719. https://doi.org/10.1109/ACCESS.2026.3662415
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.