Understanding performance of distributed data-intensive applications

3Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

Grids, clouds and cloud-like infrastructures are capable of supporting a broad range of data-intensive applications. There are interesting and unique performance issues that appear as the volume of data and degree of distribution increases. New scalable data-placement and management techniques, as well as novel approaches to determine the relative placement of data and computational workload, are required. We develop and study a genome sequence matching application that is simple to control and deploy, yet serves as a prototype of a data-intensive application. The application uses a SAGA-based implementation of the All-Pairs pattern. This paper aims to understand some of the factors that influence the performance of this application and the interplay of those factors. We also demonstrate how the SAGA approach can enable data-intensive applications to be extensible and interoperable over a range of infrastructure. This capability enables us to compare and contrast two different approaches for executing distributed data-intensive applications-simple application-level data-placement heuristics versus distributed file systems. © 2010 The Royal Society.

Cite

CITATION STYLE

APA

Miceli, C., Miceli, M., Rodriguez-Milla, B., & Jha, S. (2010). Understanding performance of distributed data-intensive applications. In Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences (Vol. 368, pp. 4089–4102). Royal Society. https://doi.org/10.1098/rsta.2010.0168

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free