Sampling Techniques to Overcome Class Imbalance in a Cyberbullying Context

  • Colton D
  • Hofmann M
N/ACitations
Citations of this article
15Readers
Mendeley users who have this article in their library.

Abstract

The majority of datasets suffer from class imbalance where samples of a dominant class significantly outnumber the samples available for the minority class that is to be detected. Prediction and classification machine learning models work best when there are roughly equal numbers of each class type. This paper explores sampling techniques that can be used to overcome this class imbalance problem in a cyberbullying context. A newly classified cyberbullying dataset, including detailed descriptions of the criteria used in its classification, was used to examine the feasibility of applying text mining techniques, to automate the detection of cyberbullying text when the dataset shows a significant class imbalance between the positive, cyberbullying, sample and the negative, not cyberbullying, samples. In this paper, we will investigate if oversampling the minority positive class or undersampling the majority negative class affects the performance of a prediction model. A compromise solution where the positive class is partially oversampled, and the negative class is partially undersampled is also examined. Although not strictly a class imbalance solution, sampling using the most frequently observed features was also explored.

Cite

CITATION STYLE

APA

Colton, D., & Hofmann, M. (2019). Sampling Techniques to Overcome Class Imbalance in a Cyberbullying Context. Journal of Computer-Assisted Linguistic Research, 3(1), 21. https://doi.org/10.4995/jclr.2019.11112

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free