Pengklasifikasian Dokumen Berbahasa Indonesia Dengan Pengindeksan Berbasis LSI

  • Ridok A
  • . I
N/ACitations
Citations of this article
26Readers
Mendeley users who have this article in their library.

Abstract

Classification of text documents aimed to determine the category of a document based on its similarity to set of documents which have been previously labeled. However, most existing methods of classification were conducted based on key words or words that are considered important by assuming each representing a unique concept. Whereas in fact some of the words that have the same meaning or semantics should be represented as a unique word. In this research LSI -based approach used on KNN to classify documents in Indonesian language. Weighting the terms of the training documents or testing using tf-idf, which represented respectively in term-document matrix A and B. Furthermore, the matrix A is decomposed using SVD to obtain matrices U and V are reduced by k-rank. Both matrices U and V are used to reduce B as a representation of test documents. The best system performance evaluation based on the results obtained LSI-based in the KNN classification without stemming with threshould 2. However, the best performance evaluation based on the time achieved when KNN LSI with stemming the KNN with threshould 5. Performance-based LSI is significantly much better than the tradisional KNN in term both the outcome and timing.

Cite

CITATION STYLE

APA

Ridok, A., & . I. (2015). Pengklasifikasian Dokumen Berbahasa Indonesia Dengan Pengindeksan Berbasis LSI. Jurnal Teknologi Informasi Dan Ilmu Komputer, 2(2), 87. https://doi.org/10.25126/jtiik.201522136

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free