Language Identification and Analysis of Code-Switched Social Media Text

N/ACitations
Citations of this article
113Readers
Mendeley users who have this article in their library.

Abstract

In this paper, we detail our work on comparing different word-level language identification systems for code-switched Hindi-English data and a standard Spanish-English dataset. In this regard, we build a new code-switched dataset for Hindi-English. To understand the code-switching patterns in these language pairs, we investigate different codeswitching metrics. We find that the CRF model outperforms the neural network based models by a margin of 2-5 percentage points for Spanish-English and 3-5 percentage points for Hindi-English.

Cite

CITATION STYLE

APA

Mave, D., Maharjan, S., & Solorio, T. (2018). Language Identification and Analysis of Code-Switched Social Media Text. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 51–61). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w18-3206

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free