Abstract
Text clustering is a data mining technique that is becoming more important in present studies. Document clustering makes use of text clustering to divide documents according to the various topics. The choice of words in document clustering is important to ensure that the document can be classified correctly. Three different methods of clustering which are hierarchical clustering, k-means and k-medoids are used and compared in this study in order to identify the best method which produce the best result in document clustering. The three methods are applied on 60 sports articles involving four different types of sports. The k-medoids clustering produced the worst result while k-means clustering is found to be more sensitive towards general words. Therefore, the method of hierarchical clustering is deemed more stable to produce a meaningful result in document clustering analysis.
Author supplied keywords
Cite
CITATION STYLE
Ariff, N. M., Bakar, M. A. A., & Rahmad, M. I. (2018). Comparative study of document clustering algorithms. International Journal of Engineering and Technology(UAE), 7(4), 246–251. https://doi.org/10.14419/ijet.v7i4.11.20816
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.