Under-Sampling Approaches for Improving Prediction of the Minority Class in an Imbalanced Dataset

Show Jane Yen; Yue Shi Lee

Book Chapter

Under-Sampling Approaches for Improving Prediction of the Minority Class in an Imbalanced Dataset

Springer Science and Business Media Deutschland GmbH, (2006), 731-740

DOI: 10.1007/978-3-540-37256-1_89

13Citations

71Readers

Get full text

Abstract

The most important factor of classification for improving classification accuracy is the training data. However, the data in real-world applications often are imbalanced class distribution, that is, most of the data are in majority class and little data are in minority class. In this case, if all the data are used to be the training data, the classifier tends to predict that most of the incoming data belong to the majority class. Hence, it is important to select the suitable training data for classification in the imbalanced class distribution problem. In this paper, we propose cluster-based under-sampling approaches for selecting the representative data as training data to improve the classification accuracy for minority class in the imbalanced class distribution problem. The experimental results show that our cluster-based under-sampling approaches outperform the other under-sampling techniques in the previous studies.

Author supplied keywords

Cite

CITATION STYLE

APA

Yen, S. J., & Lee, Y. S. (2006). Under-Sampling Approaches for Improving Prediction of the Minority Class in an Imbalanced Dataset. In Lecture Notes in Control and Information Sciences (Vol. 344, pp. 731–740). Springer Science and Business Media Deutschland GmbH. https://doi.org/10.1007/978-3-540-37256-1_89

Under-Sampling Approaches for Improving Prediction of the Minority Class in an Imbalanced Dataset

Abstract

Author supplied keywords

Cite

Register to see more suggestions