The use of deep learning distributed representations in the identification of abusive text

24Citations
Citations of this article
32Readers
Mendeley users who have this article in their library.

Abstract

The selection of optimal feature representations is a critical step in the use of machine learning in text classification. Traditional features (e.g. bag of words and n-grams) have dominated for decades, but in the past five years, the use of learned distributed representations has become increasingly common. In this paper, we summarise and present a categorisation of the state-of-the-art distributed representation techniques, including word and sentence embedding models. We carry out an empirical analysis of the performance of the various feature representations using the scenario of detecting abusive comments. We compare classification accuracies across a range of off-the-shelf embedding models using 10 labelled datasets gathered from different social media platforms. Our results show that multi-task sentence embedding models perform best with consistently highest classification results in comparison to other embedding models. We hope our work can be a guideline for practitioners in selecting appropriate features in text classification task, particularly in the domain of abuse detection.

Cite

CITATION STYLE

APA

Chen, H., McKeever, S., & Delany, S. J. (2019). The use of deep learning distributed representations in the identification of abusive text. In Proceedings of the 13th International Conference on Web and Social Media, ICWSM 2019 (Vol. 13, pp. 125–133). Association for the Advancement of Artificial Intelligence. https://doi.org/10.1609/icwsm.v13i01.3215

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free