Classification in imbalanced datasets

  • Debray T
N/ACitations
Citations of this article
38Readers
Mendeley users who have this article in their library.

Abstract

AbstractIn this thesis we study the classification task in the presence of class imbal-anced data. This task arises in many applications when we are interested inthe under-represented (minority) classes. Examples of such applications arerelated to fraud detection, medical diagnosis and monitoring, text categoriza-tion, risk management, information retrieval and filtering. Although there existmany standard approaches to the classification task, most of them have poorgeneralisation performance on the minority class.This thesis studies well-known approaches to the classification problem inthe presence of class imbalanced data, such as Cost-Sensitivity, Bagging for Im-balanced Datasets, MetaCost and SMOTE. The main contribution of the thesisis a new approach to the problem that we call Naive Bayes Sampling. Theapproach is a generative approach. It generates new instances of the minorityclass by bootstrapping values of each feature present in the training data. Ex-periments show the superiority of our approach on 4 UCI datasets and a medicaldataset provided by KULeuven

Cite

CITATION STYLE

APA

Debray, T. (2009). Classification in imbalanced datasets. Faculty of Humanities and Sciences, Maastricht University, 4(3), 78–82. Retrieved from http://www.mkbgoogle.com/public/papers/MScThesis_ClassImbalance.pdf

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free