UniParma at SemEval-2021 Task 5: Toxic Spans Detection Using CharacterBERT and Bag-of-Words Model

3Citations
Citations of this article
48Readers
Mendeley users who have this article in their library.

Abstract

With the ever-increasing availability of digital information, toxic content is also on the rise. Therefore, the detection of this type of language is of paramount importance. We tackle this problem utilizing a combination of a state-of-the-art pre-trained language model (CharacterBERT) and a traditional bag-of-words technique. Since the content is full of toxic words that have not been written according to their dictionary spelling, attendance to individual characters is crucial. Therefore, we use CharacterBERT to extract features based on the word characters. It consists of a Character-CNN module that learns character embeddings from the context. These are, then, fed into the well-known BERT architecture. The bag-of-words method, on the other hand, further improves upon that by making sure that some frequently used toxic words get labeled accordingly. With a ∼4 percent difference from the first team, our system ranked 36th in the competition. The code is available for further research and reproduction of the results1

References Powered by Scopus

BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant Supervision

203Citations
N/AReaders
Get full text

Cited by Powered by Scopus

Leveraging fusion of sequence tagging models for toxic spans detection

7Citations
N/AReaders
Get full text

SOLD: Sinhala offensive language dataset

3Citations
N/AReaders
Get full text

Evaluation of Word Embeddings for Toxic Span Prediction

0Citations
N/AReaders
Get full text

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Cite

CITATION STYLE

APA

Karimi, A., Rossi, L., & Prati, A. (2021). UniParma at SemEval-2021 Task 5: Toxic Spans Detection Using CharacterBERT and Bag-of-Words Model. In SemEval 2021 - 15th International Workshop on Semantic Evaluation, Proceedings of the Workshop (pp. 220–224). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2021.semeval-1.25

Readers' Seniority

Tooltip

PhD / Post grad / Masters / Doc 9

60%

Researcher 4

27%

Professor / Associate Prof. 1

7%

Lecturer / Post doc 1

7%

Readers' Discipline

Tooltip

Computer Science 14

70%

Linguistics 4

20%

Neuroscience 1

5%

Social Sciences 1

5%

Save time finding and organizing research with Mendeley

Sign up for free