Abstract
Author attribution is the problem of assigning an author to an unknown text. We propose a new approach to solve such a problem using an extended version of the probabilistic context free grammar language model, supplied by more informative lexical and syntactic features. In addition to the probabilities of the production rules in the generated model, we add probabilities to terminals, non-terminals, and punctuation marks. Also, the new model is augmented with a scoring function which assigns a score for each production rule. Since the new model contains different features, optimum weights, found using a genetic algorithm, are added to the model to govern how each feature participates in the classification. The advantage of using many features is to successfully capture the different writing styles of authors. Also, using a scoring function identifies the most discriminative rules. Using optimum weights supports capturing different authors' styles, which increases the classifier's performance. The new model is tested over nine authors, 20 Arabic documents per author, where the training and testing are done using the leave-one-out method. The initial error rate of the system is 20. 6%. Using the optimum weights for features reduces the error rate to 12.8%.
Cite
CITATION STYLE
Abuhaiba, I. S. I., & Eltibi, M. F. (2016). Author attribution of arabic texts using extended probabilistic context free grammar language model. International Journal of Intelligent Systems and Applications, 8(6), 27–39. https://doi.org/10.5815/ijisa.2016.06.04
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.