Multi-label Masked Language Modeling on Zero-shot Code-switched Sentiment Analysis

1Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.
Get full text

Abstract

In multilingual communities, code-switching is a common phenomenon and code-switched tasks have become a crucial area of research in natural language processing (NLP) applications. Existing approaches mainly focus on supervised learning. However, it is expensive to annotate a sufficient amount of code-switched data. In this paper, we consider zero-shot setting and improve model performance on code-switched tasks via monolingual language datasets, unlabeled code-switched datasets, and semantic dictionaries. Inspired by the mechanism of code-switching itself, we propose multi-label masked language modeling and predict both the masked word and its synonyms in other languages. Experimental results show that compared with baselines, our method can further improve the pretrained multilingual model's performance on code-switched sentiment analysis datasets.

Cite

CITATION STYLE

APA

Li, Z., Gao, X., Zhang, J., & Zhang, Y. (2022). Multi-label Masked Language Modeling on Zero-shot Code-switched Sentiment Analysis. In SIGIR 2022 - Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 2663–2668). Association for Computing Machinery, Inc. https://doi.org/10.1145/3477495.3531914

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free