Reliable Actors with Retry Orchestration

5Citations
Citations of this article
16Readers
Mendeley users who have this article in their library.

Abstract

Cloud developers have to build applications that are resilient to failures and interruptions. We advocate for a fault-Tolerant programming model for the cloud based on actors, retry orchestration, and tail calls. This model builds upon persistent data stores and message queues readily available on the cloud. Retry orchestration not only guarantees that (1) failed actor invocations will be retried but also that (2) completed invocations are never repeated and (3) it preserves a strict happen-before relationship across failures within call stacks. Tail calls can break complex tasks into simple steps to minimize re-execution during recovery. We review key application patterns and failure scenarios. We formalize a process calculus to precisely capture the mechanisms of fault tolerance in this model. We briefly describe our implementation. Using an application inspired by a typical enterprise scenario, we validate the functional correctness of our implementation and assess the impact of fault preparedness and recovery on performance.

Cite

CITATION STYLE

APA

Tardieu, O., Grove, D., Bercea, G. T., Castro, P., Cwiklik, J., & Epstein, E. (2023). Reliable Actors with Retry Orchestration. Proceedings of the ACM on Programming Languages, 7, 1293–1316. https://doi.org/10.1145/3591273

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free