Named Entity Recognition plays a major role in several downstream applications in NLP. Though this task has been heavily studied in formal monolingual texts and also noisy texts like Twitter data, it is still an emerging task in code-switched (CS) content on social media. This paper describes our participation in the shared task of NER on code-switched data for Spanglish (Spanish + English) and Arabish (Arabic + English). In this paper we describe models that intuitively developed from the data for the shared task Named Entity Recognition on Code-switched Data. Owing to the sparse and non-linear relationships between words in Twitter data, we explored neural architectures that are capable of non-linearities fairly well. In specific, we trained character level models and word level models based on Bidirectional LSTMs (Bi-LSTMs) to perform sequential tagging. We trained multiple models to identify nominal mentions and subsequently used this information to predict the labels of named entity in a sequence. Our best model is a character level model along with word level pre-trained multilingual embeddings that gave an F-score of 56.72 in Spanglish and a word level model that gave an F-score of 65.02 in Arabish on the test data.
CITATION STYLE
Geetha, P., Chandu, K. R., & Black, A. W. (2018). Tackling Code-Switched NER: Participation of CMU. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 126–131). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w18-3217
Mendeley helps you to discover research relevant for your work.