Abstract
The typical representative of the hard clustering algorithm, K-means, is one of the fastest processing algorithms with good scalability. However, it cannot deal with categorical attributes, which is one of the important indicators to measure the pros and cons. Due to the lack of processing capabilities on categorical attributes, k-means has a large limit on data processing capabilities. This paper proposes a clustering algorithm extends from K-means. This algorithm introduces the concept of a Pseudo-mean distance calculation formula and a counting-table so that categorical attributes can be processed while reducing the time cost as much as possible. Experimental results illustrate the proposed Pseudo-means can extend the processing range of k-means to category-type data, and the counting table also effectively reduces the time cost.
Cite
CITATION STYLE
Xi, C., & Shuo, P. (2020). Efficient EK-means: Extended K-means Clustering for Categorical data with High Processing Speed. In Journal of Physics: Conference Series (Vol. 1584). Institute of Physics Publishing. https://doi.org/10.1088/1742-6596/1584/1/012074
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.