Improving Arabic Text Categorization using Normalization and Stemming Techniques

  • M. R
  • M. H
  • Hussein M
N/ACitations
Citations of this article
52Readers
Mendeley users who have this article in their library.

Abstract

Text Categorization is a technique for assigning documents based on their contents to one or more pre-defined categories. Achieving highest categorization accuracy remains one of the major challenges and it is also time consuming. We proposed approach to tackle these challenges. The proposed approach uses Frequency Ratio Accumulation Method (FRAM) as a classifier. Its features are represented using bag of word technique and an improved Term Frequency (TF) technique is used in features selection. The proposed approach is tested with known datasets. The experiments are done without both of normalization and stemming, with one of them, and with both of them. The obtained results of proposed approach are generally improved compared to existing techniques.The performance attributes of proposed Arabic Text Categorization approach were considered: Accuracy, Recall, Precision and F-measure (F1). The averages of the obtained results are 97.50%, 97.50%, 97.51%, and 97.49% respectively using normalization. Keywords Arabic text categorization, Frequency ratio accumulation method (FRAM), Bag-Of-Word (BOW), Features selection, Term and document frequency.

Cite

CITATION STYLE

APA

M., R., M., H., & Hussein, M. (2016). Improving Arabic Text Categorization using Normalization and Stemming Techniques. International Journal of Computer Applications, 135(2), 38–43. https://doi.org/10.5120/ijca2016908328

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free