DRASH: A data replication-aware scheduler in geo-distributed data centers

18Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Driven by the trends of BigData and Cloud computing, there is a growing demand for processing and analyzing data that are generated and stored across geo-distributed data centers. However, due to the limited network bandwidth between data centers and the growing data volume spread across different locations, it has become increasingly inefficient to aggregate data and to perform computations at a single data center. An approach that has been commonly used by data-intensive cluster computation systems, like Hadoop, is to distribute computations based on data locality so that data can be processed locally to reduce the network overhead and improve performance. But limited work has been done to adapt and evaluate such technique for geo-distributed data centers. In this paper, we proposed DRASH (Data-Replication Aware Scheduler), a job scheduling algorithm that enforces data locality to prevent data transfer, and exploits data replications to improve overall system performance. Our evaluation using simulations with realistic workload traces shows that DRASH can outperform other existing approaches by 16% to 60% in average job completion time, and achieve greater improvements under higher data replication factors.

Cite

CITATION STYLE

APA

Convolbo, M. W., Chou, J., Lu, S., & Chung, Y. C. (2017). DRASH: A data replication-aware scheduler in geo-distributed data centers. In Proceedings of the International Conference on Cloud Computing Technology and Science, CloudCom (Vol. 0, pp. 302–309). IEEE Computer Society. https://doi.org/10.1109/CloudCom.2016.0056

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free