Acorn: Aggressive Result Caching in Distributed Data Processing Frameworks

7Citations
Citations of this article
11Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Result caching is crucial to the performance of data processing systems, but two trends complicate its use. First, immutable datasets make it difficult to efficiently employ powerful result caching techniques like predicate analysis, since predicate analysis typically requires optimized query plans but generating those plans can be costly with data immutability. Second, increased support for user-defined functions (UDFs), which are treated as black boxes by query engines, hinders aggressive result caching. This paper overcomes these problems by introducing 1) a judicious adaptation of predicate analysis on analyzed query plans that avoids unnecessary query optimization, and 2) a UDF translator that transparently compiles UDFs from general purpose languages into native equivalents. We then present Acorn, a concrete implementation of these techniques in Spark SQL that provides speedups of up to 5x across multiple benchmark and real Spark graph processing workloads.

Cite

CITATION STYLE

APA

Ramjit, L., Interlandi, M., Wu, E., & Netravali, R. (2019). Acorn: Aggressive Result Caching in Distributed Data Processing Frameworks. In SoCC 2019 - Proceedings of the ACM Symposium on Cloud Computing (pp. 206–219). Association for Computing Machinery. https://doi.org/10.1145/3357223.3362702

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free