Abstract
Imbalanced data is common and presents significant challenge towards classification of data. In this research, we present a combination of two techniques used for handling class imbalance in datasets, SMOTE (Synthetic Minority Over-sampling Technique) and Tomek Links. Each strategy handles the class imbalance problem in a unique way, and their combination attempts to create a more balanced and cleaner dataset for training machine learning models to handle binary classification by addressing problematic or difficult-to-classify data. Machine learning classifiers used in this study are K-Nearest Neighbour (KNN), Support Vector Machine (SVM), Logistic Regression, Decision Tree (DT), Random Forest (RF), Gradient Boosting, Extreme Gradient Boosting (XGBoost), Light Gradient Boosting (LGBM), AdaBoost and Catboost. It has been discovered that the mean F1 score for resampled datasets provides more trustworthy results for forecasting floods.
Author supplied keywords
Cite
CITATION STYLE
Zuhairi, A. H., Yakub, F., Omar, M., Sharifuddin, M., Razak, K. A., & Faruq, A. (2024). Imbalanced Flood Forecast Dataset Resampling Using SMOTE-Tomek Link. In International Exchange and Innovation Conference on Engineering and Sciences (Vol. 10, pp. 845–850). Kyushu University. https://doi.org/10.5109/7323359
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.