Comparison of distributed K-means and distributed fuzzy C-means algorithms for text clustering

3Citations
Citations of this article
14Readers
Mendeley users who have this article in their library.

Abstract

Text clustering has been developed in distributed system due to increasing data. The popular algorithms like K-Means (KM) and Fuzzy C-Means (FCM) are combined with Map Reduce algorithm in Hadoop Environment to be distributable and parallelizable. The problem is performance comparison between Distributed KM (DKM) and Distributed FCM (DFCM) that uses Tanimoto Distance Measure (TDM) has not been studied yet. It is important because TDM-s characteristics are scale invariant while allowing discrimination collinear vectors. This work compared the combination of TDM with DKM (DKM-T) and TDM with DFCM (DFCM-T) to acquire performance of both algorithms. The result shows that DFCM-T has better intra-cluster and inter-cluster densities than those of DKM-T. Moreover, DFCM-T has lower processing time than that of DKM-T when total nodes used are 4 and 8. DFCM-T and DKM-T can perform clustering of 1,400,000 text files in 16.18 and 9.74 minutes but the preprocessing times take hours to complete.

Cite

CITATION STYLE

APA

Made Artha Agastya, I., Adji, T. B., & Setiawan, N. A. (2017). Comparison of distributed K-means and distributed fuzzy C-means algorithms for text clustering. Communications in Science and Technology, 2(1), 11–17. https://doi.org/10.21924/cst.2.1.2017.46

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free