Text Document Classification: An Approach Based on Indexing

  • Harish B
N/ACitations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

In this paper we propose a new method of classifying text documents. Unlike conventional vector space models, the proposed method preserves the sequence of term occurrence in a document. The term sequence is effectively preserved with the help of a novel datastructure called 'Status Matrix'. Further the corresponding classification technique has been proposed for efficient classification of text documents. In addition, in order to avoid sequential matching during classification, we propose to index the terms in B-tree, an efficient index scheme. Each term in B-tree is associated with a list of class labels of those documents which contain the term. Further the corresponding classification technique has been proposed. To corroborate the efficacy of the proposed representation and status matrix based classification, we have conducted extensive experiments on various datasets.

Cite

CITATION STYLE

APA

Harish, B. S. (2012). Text Document Classification: An Approach Based on Indexing. International Journal of Data Mining & Knowledge Management Process, 2(1), 43–62. https://doi.org/10.5121/ijdkp.2012.2104

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free