Event Annotation and Detection in Kannada-English Code-Mixed Social Media Data

S. Sumukh; Abhinav Appidi; Manish Shrivastava

Conference ProceedingsOPEN ACCESS

Event Annotation and Detection in Kannada-English Code-Mixed Social Media Data

International Conference Recent Advances in Natural Language Processing, RANLP (2023) 1007-1014

DOI: 10.26615/978-954-452-092-2_108

0Citations

6Readers

Abstract

Code-mixing (CM) is a frequently observed phenomenon on social media platforms in multilingual societies such as India. While the increase in code-mixed content on these platforms provides good amount of data for studying various aspects of code-mixing, the lack of automated text analysis tools makes such studies difficult. To overcome the same, tools such as language identifiers, Parts-of-Speech (POS) taggers and Named Entity Recognition (NER) for analysing code-mixed data have been developed. One such important tool is Event Detection, an important information retrieval task which can be used to identify critical facts occurring in the vast streams of unstructured text data available. While event detection from text is a hard problem on its own, social media data adds to it with its informal nature, and codemixed (Kannada-English) data further complicates the problem due to its word-level mixing, lack of structure and incomplete information. In this work, we have tried to address this problem. We have proposed guidelines for the annotation of events in Kannada-English CM data and provided some baselines for the same with careful feature selection.

Cite

CITATION STYLE

APA

Sumukh, S., Appidi, A., & Shrivastava, M. (2023). Event Annotation and Detection in Kannada-English Code-Mixed Social Media Data. In International Conference Recent Advances in Natural Language Processing, RANLP (pp. 1007–1014). Incoma Ltd. https://doi.org/10.26615/978-954-452-092-2_108

Event Annotation and Detection in Kannada-English Code-Mixed Social Media Data

Abstract

Cite

Register to see more suggestions