MS@IW at SemEval-2022 Task 4: Patronising and Condescending Language Detection with Synthetically Generated Data

1Citations
Citations of this article
27Readers
Mendeley users who have this article in their library.

Abstract

In this description paper we outline the system architecture submitted to Task 4, Subtask 1 at SemEval-2022. We leverage the generative power of state-of-the-art generative pretrained transformer models to increase training set size and remedy class imbalance issues. Our best submitted system is trained on a synthetically enhanced dataset with 10.3 times as many positive samples as the original dataset and reaches an F1 score of 50.62%, which is 10 percentage points higher than our initial system trained on an undersampled version of the original dataset. We explore possible reasons for the comparably low score in the overall task ranking and report on experiments conducted during the post-evaluation phase.

Cite

CITATION STYLE

APA

Meyer, S., Schmidhuber, M., & Kruschwitz, U. (2022). MS@IW at SemEval-2022 Task 4: Patronising and Condescending Language Detection with Synthetically Generated Data. In SemEval 2022 - 16th International Workshop on Semantic Evaluation, Proceedings of the Workshop (pp. 363–368). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.semeval-1.47

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free