Comparative study of document clustering algorithms

N. M. Ariff; M. A.A. Bakar; M. I. Rahmad

Journal ArticleOPEN ACCESS

Comparative study of document clustering algorithms

International Journal of Engineering and Technology(UAE) (2018) 7(4) 246-251

DOI: 10.14419/ijet.v7i4.11.20816

5Citations

31Readers

Abstract

Text clustering is a data mining technique that is becoming more important in present studies. Document clustering makes use of text clustering to divide documents according to the various topics. The choice of words in document clustering is important to ensure that the document can be classified correctly. Three different methods of clustering which are hierarchical clustering, k-means and k-medoids are used and compared in this study in order to identify the best method which produce the best result in document clustering. The three methods are applied on 60 sports articles involving four different types of sports. The k-medoids clustering produced the worst result while k-means clustering is found to be more sensitive towards general words. Therefore, the method of hierarchical clustering is deemed more stable to produce a meaningful result in document clustering analysis.

Author supplied keywords

Cite

CITATION STYLE

APA

Ariff, N. M., Bakar, M. A. A., & Rahmad, M. I. (2018). Comparative study of document clustering algorithms. International Journal of Engineering and Technology(UAE), 7(4), 246–251. https://doi.org/10.14419/ijet.v7i4.11.20816

Comparative study of document clustering algorithms

Abstract

Author supplied keywords

Cite

Register to see more suggestions