Comparative study of document clustering algorithms

5Citations
Citations of this article
31Readers
Mendeley users who have this article in their library.

Abstract

Text clustering is a data mining technique that is becoming more important in present studies. Document clustering makes use of text clustering to divide documents according to the various topics. The choice of words in document clustering is important to ensure that the document can be classified correctly. Three different methods of clustering which are hierarchical clustering, k-means and k-medoids are used and compared in this study in order to identify the best method which produce the best result in document clustering. The three methods are applied on 60 sports articles involving four different types of sports. The k-medoids clustering produced the worst result while k-means clustering is found to be more sensitive towards general words. Therefore, the method of hierarchical clustering is deemed more stable to produce a meaningful result in document clustering analysis.

Cite

CITATION STYLE

APA

Ariff, N. M., Bakar, M. A. A., & Rahmad, M. I. (2018). Comparative study of document clustering algorithms. International Journal of Engineering and Technology(UAE), 7(4), 246–251. https://doi.org/10.14419/ijet.v7i4.11.20816

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free