Abstract
MapReduce is a suitable and efficient parallel programming pattern for processing big data analysis. In recent years, many frameworks/languages have implemented this pattern to achieve high performance in data mining applications, particularly for distributed memory architectures (e.g., clusters). Nevertheless, the industry of processors is now able to offer powerful processing on single machines (e.g., multi-core). Thus, these applications may address the parallelism in another architectural level. The target problems of this paper are code reuse and programming effort reduction since current solutions do not provide a single interface to deal with these two architectural levels. Therefore, we propose a unified domain-specific language in conjunction with transformation rules for code generation for Hadoop and Phoenix++. We selected these frameworks as state-of-the-art MapReduce implementations for distributed and shared memory architectures, respectively. Our solution achieves a programming effort reduction from 41.84% and up to 95.43% without significant performance losses (below the threshold of 3%) compared to Hadoop and Phoenix++.
Author supplied keywords
Cite
CITATION STYLE
Adornes, D., Griebler, D., Ledur, C., & Fernandes, L. G. (2015). A unified MapReduce domain-specific language for distributed and shared memory architectures. In Proceedings of the International Conference on Software Engineering and Knowledge Engineering, SEKE (Vol. 2015-January, pp. 619–624). Knowledge Systems Institute Graduate School. https://doi.org/10.18293/SEKE2015-204
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.