Clustering and selecting categorical features

Cláudia Silvestre; Margarida Cardoso; Mário Figueiredo

Conference Proceedings

Clustering and selecting categorical features

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2013) 8154 LNAI 331-342

DOI: 10.1007/978-3-642-40669-0_29

0Citations

5Readers

Get full text

Abstract

In data clustering, the problem of selecting the subset of most relevant features from the data has been an active research topic. Feature selection for clustering is a challenging task due to the absence of class labels for guiding the search for relevant features. Most methods proposed for this goal are focused on numerical data. In this work, we propose an approach for clustering and selecting categorical features simultaneously. We assume that the data originate from a finite mixture of multinomial distributions and implement an integrated expectation-maximization (EM) algorithm that estimates all the parameters of the model and selects the subset of relevant features simultaneously. The results obtained on synthetic data illustrate the performance of the proposed approach. An application to real data, referred to official statistics, shows its usefulness. © 2013 Springer-Verlag.

Author supplied keywords

Cite

CITATION STYLE

APA

Silvestre, C., Cardoso, M., & Figueiredo, M. (2013). Clustering and selecting categorical features. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 8154 LNAI, pp. 331–342). https://doi.org/10.1007/978-3-642-40669-0_29

Clustering and selecting categorical features

Abstract

Author supplied keywords

Cite

Register to see more suggestions