Abstract
Checkpoint-rollback recovery, which is a universal method for restoring distributed systems after faults, requires a sophisticated snapshot algorithm especially if the systems are large-scale, since repeatedly taking global snapshots of the whole system requires unacceptable communication cost. As a sophisticated snapshot algorithm, a partial snapshot algorithm has been introduced that takes a snapshot of a subsystem consisting only of the nodes that are communication-related to the initiator instead of a global snapshot of the whole system. In this paper, we modify the previous partial snapshot algorithm to create a new one that can take a partial snapshot more efficiently, especially when multiple nodes concurrently initiate the algorithm. Experiments show that the proposed algorithm greatly reduces the amount of communication needed for taking partial snapshots. Copyright © 2014 The Institute of Electronics, Inf rmation and Communication Engineers.
Author supplied keywords
Cite
CITATION STYLE
Kim, Y., Araragi, T., Nakamura, J., & Masuzawa, T. (2014). A concurrent partial snapshot algorithm for large-scale and dynamic distributed systems. IEICE Transactions on Information and Systems, E97-D(1), 65–76. https://doi.org/10.1587/transinf.E97.D.65
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.