Abstract
In this paper, we detail our work on comparing different word-level language identification systems for code-switched Hindi-English data and a standard Spanish-English dataset. In this regard, we build a new code-switched dataset for Hindi-English. To understand the code-switching patterns in these language pairs, we investigate different codeswitching metrics. We find that the CRF model outperforms the neural network based models by a margin of 2-5 percentage points for Spanish-English and 3-5 percentage points for Hindi-English.
Cite
CITATION STYLE
Mave, D., Maharjan, S., & Solorio, T. (2018). Language Identification and Analysis of Code-Switched Social Media Text. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 51–61). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w18-3206
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.