A concurrent partial snapshot algorithm for large-scale and dynamic distributed systems

5Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

Checkpoint-rollback recovery, which is a universal method for restoring distributed systems after faults, requires a sophisticated snapshot algorithm especially if the systems are large-scale, since repeatedly taking global snapshots of the whole system requires unacceptable communication cost. As a sophisticated snapshot algorithm, a partial snapshot algorithm has been introduced that takes a snapshot of a subsystem consisting only of the nodes that are communication-related to the initiator instead of a global snapshot of the whole system. In this paper, we modify the previous partial snapshot algorithm to create a new one that can take a partial snapshot more efficiently, especially when multiple nodes concurrently initiate the algorithm. Experiments show that the proposed algorithm greatly reduces the amount of communication needed for taking partial snapshots. Copyright © 2014 The Institute of Electronics, Inf rmation and Communication Engineers.

Cite

CITATION STYLE

APA

Kim, Y., Araragi, T., Nakamura, J., & Masuzawa, T. (2014). A concurrent partial snapshot algorithm for large-scale and dynamic distributed systems. IEICE Transactions on Information and Systems, E97-D(1), 65–76. https://doi.org/10.1587/transinf.E97.D.65

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free