Efficient Multilingual Text Classification for Indian Languages

10Citations
Citations of this article
47Readers
Mendeley users who have this article in their library.
Get full text

Abstract

India is one of the richest language hubs on the earth and is very diverse and multilingual. But apart from a few Indian languages, most of them are still considered to be resource poor. Since most of the NLP techniques either require linguistic knowledge that can only be developed by experts and native speakers of that language or they require a lot of labelled data which is again expensive to generate, the task of text classification becomes challenging for most of the Indian languages. The main objective of this paper is to see how one can benefit from the lexical similarity found in Indian languages in a multilingual scenario. Can a classification model trained on one Indian language be reused for other Indian languages? So, we performed zero-shot text classification via exploiting lexical similarity and we observed that our model performs best in those cases where the vocabulary overlap between the language datasets is maximum. Our experiments also confirm that a single multilingual model trained via exploiting language relatedness outperforms the baselines by significant margins.

Cite

CITATION STYLE

APA

Aggarwal, S., Kumar, S., & Mamidi, R. (2021). Efficient Multilingual Text Classification for Indian Languages. In International Conference Recent Advances in Natural Language Processing, RANLP (pp. 19–25). Incoma Ltd. https://doi.org/10.26615/978-954-452-072-4_003

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free