Samawa Language: Part of Speech Tagset and Tagged Corpus for NLP Resources

1Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Samawa language is one of the local languages in Sumbawa Island, Indonesia. It was categorized as an under-resourced language since there is only a small number of books written and studies reported in the literature. The availability of Samawa corpus is an alternative to preserve Samawa as a cultural heritage. In this work, we describe our first effort to build the first Samawa corpus and assign 24 part of speech information to each token based on lexical and grammatical rules of Samawa. The raw data used was collected from the manuscripts, textbooks, magazines, and text from websites. It was cleaned and normalized by UTF-8 Unicode standard scheme and converted into XML format with TEI guidelines. The result is a Samawa tagged corpus of 739 sentences that contain 11,799 tokens and can be used for developing tools in many NLP applications.

Cite

CITATION STYLE

APA

Hariyanti, T., Aida, S., & Kameda, H. (2018). Samawa Language: Part of Speech Tagset and Tagged Corpus for NLP Resources. In Journal of Physics: Conference Series (Vol. 1061). Institute of Physics Publishing. https://doi.org/10.1088/1742-6596/1061/1/012007

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free