Morpheme Matching Based Text Tokenization for a Scarce Resourced Language

12Citations
Citations of this article
27Readers
Mendeley users who have this article in their library.

Abstract

Text tokenization is a fundamental pre-processing step for almost all the information processing applications. This task is nontrivial for the scarce resourced languages such as Urdu, as there is inconsistent use of space between words. In this paper a morpheme matching based approach has been proposed for Urdu text tokenization, along with some other algorithms to solve the additional issues of boundary detection of compound words, affixation, reduplication, names and abbreviations. This study resulted into 97.28% precision, 93.71% recall, and 95.46% F1-measure; while tokenizing a corpus of 57000 words by using a morpheme list with 6400 entries. © 2013 Rehman et al.

Cite

CITATION STYLE

APA

Rehman, Z., Anwar, W., Bajwa, U. I., Xuan, W., & Chaoying, Z. (2013). Morpheme Matching Based Text Tokenization for a Scarce Resourced Language. PLoS ONE, 8(8). https://doi.org/10.1371/journal.pone.0068178

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free