Adaptive Distributed Streaming Similarity Joins

2Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.
Get full text

Abstract

How can we perform similarity joins of multi-dimensional streams in a distributed fashion, achieving low latency? Can we adaptively repartition those streams in order to retain high performance under concept drifts? Current approaches to similarity joins are either restricted to single-node deployments or focus on set-similarity joins, failing to cover the ubiquitous case of metric-space similarity joins. In this paper, we propose the first adaptive distributed streaming similarity join approach that gracefully scales with variable velocity and distribution of multi-dimensional data streams. Our approach can adaptively rebalance the load of nodes in the case of concept drifts, allowing for similarity computations in the general metric space. We implement our approach on top of Apache Flink and evaluate its data partitioning and load balancing schemes on a set of synthetic datasets in terms of latency, comparisons ratio, and data duplication ratio.

Cite

CITATION STYLE

APA

Siachamis, G., Psarakis, K., Fragkoulis, M., Papapetrou, O., Van Deursen, A., & Katsifodimos, A. (2023). Adaptive Distributed Streaming Similarity Joins. In DEBS 2023 - Proceedings of the 17th ACM International Conference on Distributed and Event-based Systems (pp. 25–36). Association for Computing Machinery, Inc. https://doi.org/10.1145/3583678.3596891

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free