Abstract
Clustering with heterogeneous variables in a dataset is no doubt a challenging process owing to different scales in a data. The paper introduced a SimMultiCorrData package in R to generate the artificial dataset for clustering. The construction of artificial dataset with various distribution helps to mimic the scenario of nature of real datasets. Our experiments shows that the clusterability of a dataset are influenced by various factors such as overlapping clusters, noise, sub-cluster, and unbalance objects within the clusters.
Author supplied keywords
Cite
CITATION STYLE
Shamsuddin, N. R., & Mahat, N. I. (2019). Investigation on the clusterability of heterogeneous dataset by retaining the scale of variables. Mathematics and Statistics, 7(4), 49–57. https://doi.org/10.13189/ms.2019.070707
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.