Abstract
This paper addresses the problem of sentiment analysis for Jopara, a code-switching language between Guarani and Spanish. We first collect a corpus of Guarani-dominant tweets and discuss on the difficulties of finding quality data for even relatively easy-to-annotate tasks, such as sentiment analysis. Then, we train a set of neural models, including pre-trained language models, and explore whether they perform better than traditional machine learning ones in this low-resource setup. Transformer architectures obtain the best results, despite not considering Guarani during pre-training, but traditional machine learning models perform close due to the low-resource nature of the problem.
Cite
CITATION STYLE
Agüero-Torales, M. M., Vilares, D., & López-Herrera, A. G. (2021). On the logistical difficulties and findings of Jopara Sentiment Analysis. In Computational Approaches to Linguistic Code-Switching, CALCS 2021 - Proceedings of the 5th Workshop (pp. 95–102). Association for Computational Linguistics (ACL). https://doi.org/10.26615/978-954-452-056-4_012
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.