Clustering categorical data based on combinations of attribute values

ISSN: 13494198
11Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

Clustering is an important technique for exploratory data analysis. While most of the earlier clustering algorithms focused on numerical data, real-world problems and data mining applications frequently involve categorical data. Here, we propose a new clustering algorithm for categorical data that is based on the frequency of attribute value combinations. Our algorithm finds all the combinations of attribute values in an object, which represent a subset of all the attribute values, and then groups the object using the frequency of these combinations in each cluster. As our algorithm considers all the subsets of attribute values in an object, objects in a cluster have not only similar attribute value sets but also strongly associated attribute values. Also, the proposed algorithm is not the clustering method using the similarity between only two objects, but rather uses the similarity between an object and clusters. Therefore, it provides global information in clustering results. We conducted experiments with real and synthetic data sets to evaluate FAVC. We show that FAVC is more scalable and provides higher quality results than the previous method. © 2009 ISSN.

Cite

CITATION STYLE

APA

Do, H. J., & Kim, J. Y. (2009). Clustering categorical data based on combinations of attribute values. International Journal of Innovative Computing, Information and Control, 5(12), 4393–4405.

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free