Abstract
In multilingual communities, code-switching is a common phenomenon and code-switched tasks have become a crucial area of research in natural language processing (NLP) applications. Existing approaches mainly focus on supervised learning. However, it is expensive to annotate a sufficient amount of code-switched data. In this paper, we consider zero-shot setting and improve model performance on code-switched tasks via monolingual language datasets, unlabeled code-switched datasets, and semantic dictionaries. Inspired by the mechanism of code-switching itself, we propose multi-label masked language modeling and predict both the masked word and its synonyms in other languages. Experimental results show that compared with baselines, our method can further improve the pretrained multilingual model's performance on code-switched sentiment analysis datasets.
Author supplied keywords
Cite
CITATION STYLE
Li, Z., Gao, X., Zhang, J., & Zhang, Y. (2022). Multi-label Masked Language Modeling on Zero-shot Code-switched Sentiment Analysis. In SIGIR 2022 - Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 2663–2668). Association for Computing Machinery, Inc. https://doi.org/10.1145/3477495.3531914
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.