HARALD: Augmenting Hate Speech Data Sets with Real Data

7Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.
Get full text

Abstract

The successful completion of the hate speech detection task hinges upon the availability of rich and variable labeled data, which is hard to obtain. In this work, we present a new approach for data augmentation that uses as input real unlabelled data, which is carefully selected from online platforms where invited hate speech is abundant. We show that by harvesting and processing this data (in an automatic manner), one can augment existing manually-labeled datasets to improve the classification performance of hate speech classification models. We observed an improvement in F1-score ranging from 2.7% and up to 9.5%, depending on the task (in- or cross-domain) and the model used.

Cite

CITATION STYLE

APA

Ilan, T., & Vilenchik, D. (2022). HARALD: Augmenting Hate Speech Data Sets with Real Data. In Findings of the Association for Computational Linguistics: EMNLP 2022 (pp. 2241–2248). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.findings-emnlp.335

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free