kamila: Clustering mixed-type data in R and hadoop

71Citations
Citations of this article
87Readers
Mendeley users who have this article in their library.

Abstract

In this paper we discuss the challenge of equitably combining continuous (quantita-tive) and categorical (qualitative) variables for the purpose of cluster analysis. Existing techniques require strong parametric assumptions, or difficult-to-specify tuning parameters. We describe the kamila package, which includes a weighted k-means approach to clustering mixed-type data, a method for estimating weights for mixed-type data (Modha-Spangler weighting), and an additional semiparametric method recently proposed in the literature (KAMILA). We include a discussion of strategies for estimating the number of clusters in the data, and describe the implementation of one such method in the current R package. Background and usage of these clustering methods are presented. We then show how the KAMILA algorithm can be adapted to a map-reduce framework, and implement the resulting algorithm using Hadoop for clustering very large mixed-type data sets.

Cite

CITATION STYLE

APA

Foss, A. H., & Markatou, M. (2018). kamila: Clustering mixed-type data in R and hadoop. Journal of Statistical Software, 83. https://doi.org/10.18637/jss.v083.i13

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free