An efficient distance estimation and centroid selection based on k-means clustering for small and large dataset

Girdhar Gopal Ladha; Ravi Kumar Singh Pippal

Journal ArticleOPEN ACCESS

An efficient distance estimation and centroid selection based on k-means clustering for small and large dataset

International Journal of Advanced Technology and Engineering Exploration (2020) 7(73) 234-240

DOI: 10.19101/IJATEE.2020.762109

4Citations

8Readers

Abstract

In this paper an efficient distance estimation and centroid selection based on k-means clustering for small and large dataset. Data pre-processing was performed first on the dataset. For the complete study and analysis PIMA Indian diabetes dataset was considered. After pre-processing distance and centroid estimation was performed. It includes initial selection based on randomization and then centroids updations were performed till the iterations or epochs determined. Distance measures used here are Euclidean distance (Ed), Pearson Coefficient distance (PCd), Chebyshev distance (Csd) and Canberra distance (Cad). The results indicate that all the distance algorithms performed approximately well in case of clustering but in terms of time Cad outperforms in comparison to other algorithms.

Author supplied keywords

Cite

CITATION STYLE

APA

Ladha, G. G., & Pippal, R. K. S. (2020). An efficient distance estimation and centroid selection based on k-means clustering for small and large dataset. International Journal of Advanced Technology and Engineering Exploration, 7(73), 234–240. https://doi.org/10.19101/IJATEE.2020.762109

An efficient distance estimation and centroid selection based on k-means clustering for small and large dataset

Abstract

Author supplied keywords

Cite

Register to see more suggestions