Abstract
Samawa language is one of the local languages in Sumbawa Island, Indonesia. It was categorized as an under-resourced language since there is only a small number of books written and studies reported in the literature. The availability of Samawa corpus is an alternative to preserve Samawa as a cultural heritage. In this work, we describe our first effort to build the first Samawa corpus and assign 24 part of speech information to each token based on lexical and grammatical rules of Samawa. The raw data used was collected from the manuscripts, textbooks, magazines, and text from websites. It was cleaned and normalized by UTF-8 Unicode standard scheme and converted into XML format with TEI guidelines. The result is a Samawa tagged corpus of 739 sentences that contain 11,799 tokens and can be used for developing tools in many NLP applications.
Cite
CITATION STYLE
Hariyanti, T., Aida, S., & Kameda, H. (2018). Samawa Language: Part of Speech Tagset and Tagged Corpus for NLP Resources. In Journal of Physics: Conference Series (Vol. 1061). Institute of Physics Publishing. https://doi.org/10.1088/1742-6596/1061/1/012007
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.