VACASPATI: A Diverse Corpus of Bangla Literature

N/ACitations
Citations of this article
11Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Bangla (or Bengali) is the fifth most spoken language globally; yet, the state-of-the-art NLP in Bangla is lagging for even simple tasks such as lemmatization, POS tagging, etc. This is partly due to lack of a varied quality corpus. To alleviate this need, we build VĀCASPATI, a diverse corpus of Bangla literature. The literary works are collected from various websites; only those works that are publicly available without copyright violations or restrictions are collected. We believe that published literature captures the features of a language much better than newspapers, blogs or social media posts which tend to follow only a certain literary pattern and, therefore, miss out on language variety and vocabulary. Our corpus VĀCASPATI is varied from multiple aspects, including type of composition, topic, author, time, space, etc. It contains more than 11 million sentences and 115 million words. We have also built a word embedding model, VĀC-FT, using FastText from VĀCASPATI as well as trained an Electra model, VĀC-BERT, using the corpus. VĀC-BERT has far fewer parameters and requires only a fraction of resources compared to other state-of-the-art transformer models and yet performs either better or similar on various downstream tasks. Similarly, VĀC-FT outperforms other FastText-based models on multiple downstream tasks. We also demonstrate the efficacy of VĀCASPATI as a corpus by showing that similar models built from other corpora are not as effective.

Cite

CITATION STYLE

APA

Bhattacharyya, P., Mondal, J., Maji, S., & Bhattacharya, A. (2023). VACASPATI: A Diverse Corpus of Bangla Literature. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics: Long Papers, IJCNLP-AACL 2023 (Vol. 1, pp. 1118–1130). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2023.ijcnlp-main.72

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free