Abstract
Sentiment analysis is an important task in understanding social media content like customer reviews, Twitter and Facebook feeds etc. In multilingual communities around the world, a large amount of social media text is characterized by the presence of code-switching. Thus, it has become important to build models that can handle code-switched data. However, annotated code-switched data is scarce and there is a need for unsupervised models and algorithms. We propose a general framework called Unsupervised Self-Training and show its applications for the specific use case of sentiment analysis of code-switched data. We use the power of pre-trained BERT models for initialization and fine-tune them in an unsupervised manner, only using pseudo labels produced by zero-shot transfer. We test our algorithm on multiple code-switched languages and provide a detailed analysis of the learning dynamics of the algorithm with the aim of answering the question - ‘Does our unsupervised model understand the Code-Switched languages or does it just learn its representations?’. Our unsupervised models compete well with their supervised counterparts, with their performance reaching within 1-7% (weighted F1 scores) when compared to supervised models trained for a two class problem.
Cite
CITATION STYLE
Gupta, A., Menghani, S., Rallabandi, S. K., & Black, A. W. (2021). Unsupervised Self-Training for Sentiment Analysis of Code-Switched Data. In Computational Approaches to Linguistic Code-Switching, CALCS 2021 - Proceedings of the 5th Workshop (pp. 103–112). Association for Computational Linguistics (ACL). https://doi.org/10.26615/978-954-452-056-4_013
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.