Abstract
This paper introduces the result of Team Dartmouth's experiments on each of the five subtasks for the detection of sarcasm in English and Arabic tweets. This detection was framed as a classification problem, and our contributions are threefold: we developed an English binary classifier system with RoBERTaBASE, an Arabic binary classifier with XLM-RoBERTaBASE, and an English multilabel classifier with BERTBASE. Preprocessing steps are taken with labeled input data prior to tokenization, such as extracting and appending verbs/adjectives or representative/significant keywords to the end of an input tweet to help the models better understand and generalize sarcasm detection. We also discuss the results of simple data augmentation techniques to improve the quality of the given training dataset as well as an alternative approach to the question of multilabel sequence classification. Ultimately, our systems place us in the top 14 participants for each of the five subtasks.
Cite
CITATION STYLE
Lad, R., Ma, W., & Vosoughi, S. (2022). Dartmouth at SemEval-2022 Task 6: Detection of Sarcasm. In SemEval 2022 - 16th International Workshop on Semantic Evaluation, Proceedings of the Workshop (pp. 912–918). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.semeval-1.128
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.