TNT: Text normalization based pre-training of transformers for content moderation

N/ACitations
Citations of this article
85Readers
Mendeley users who have this article in their library.

Abstract

In this work, we present a new language pre-training model TNT (Text Normalization based pre-training of Transformers) for content moderation. Inspired by the masking strategy and text normalization, TNT is developed to learn language representation by training transformers to reconstruct text from four operation types typically seen in text manipulation: substitution, transposition, deletion, and insertion. Furthermore, the normalization involves the prediction of both operation types and token labels, enabling TNT to learn from more challenging tasks than the standard task of masked word recovery. As a result, the experiments demonstrate that TNT outperforms strong baselines on the hate speech classification task. Additional text normalization experiments and case studies show that TNT is a new potential approach to misspelling correction.

Cite

CITATION STYLE

APA

Tan, F., Hu, Y., Hu, C., Li, K., & Yen, K. (2020). TNT: Text normalization based pre-training of transformers for content moderation. In EMNLP 2020 - 2020 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference (pp. 4735–4741). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.emnlp-main.383

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free