Clustering Data in Secured, Distributed Datasets

Sayantan Dey; Lee A. Carraher; Anindya Moitra; Philip A. Wilsey

Conference Proceedings

Clustering Data in Secured, Distributed Datasets

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2019) 11624 LNCS 557-572

DOI: 10.1007/978-3-030-24311-1_40

0Citations

2Readers

Get full text

Abstract

The massive growth in data generation and collection has brought to the forefront the necessity to develop mechanized methods to analyze and extract information from them. Data clustering is one of the fundamental modes to discover new insights from data. However, high dimensional data has its own challenges where many conventional clustering algorithms fails either in accuracy or scalability. To further complicate the issue, distinct subsets of sensitive data may reside in geographically separated locations with the sensitive nature of the data preventing (or inhibiting) its access for mechanized analysis. Thus, methods to discover information from the collective whole of these secured, distributed data sets that also preserves the integrity of the data must be found. In this paper we develop and assess a distributed algorithm that can cluster geographically separated data while simultaneously preserving the strict privacy requirements of non sharing of protected high dimensional data. We implement our algorithm on the distributed map-reduce based platform Spark and demonstrate its performance by comparing it to the standard data clustering algorithms.

Author supplied keywords

Cite

CITATION STYLE

APA

Dey, S., Carraher, L. A., Moitra, A., & Wilsey, P. A. (2019). Clustering Data in Secured, Distributed Datasets. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 11624 LNCS, pp. 557–572). Springer Verlag. https://doi.org/10.1007/978-3-030-24311-1_40

Clustering Data in Secured, Distributed Datasets

Abstract

Author supplied keywords

Cite

Register to see more suggestions