Network-accelerated distributed machine learning for multi-tenant settings

N/ACitations
Citations of this article
23Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Many distributed machine learning (DML) workloads are increasingly being run in shared clusters. Training in such clusters can be impeded by unexpected compute and network contention, resulting in stragglers. We present MLfabric, a contention-aware DML system that manages the performance of a DML job running in a shared cluster. The DML application hands all network communication (gradient and model transfers) to the MLfabric communication library. MLfabric then carefully orders transfers to improve convergence, opportunistically aggregates them at idle DML workers to improve resource efficiency, and replicates them to support new notions of fault tolerance, while systematically accounting for compute stragglers and network contention. We find that MLfabric achieves up to 3x speed-up in training large deep learning models in realistic dynamic cluster settings.

Cite

CITATION STYLE

APA

Viswanathan, R., Balasubramanian, A., & Akella, A. (2020). Network-accelerated distributed machine learning for multi-tenant settings. In SoCC 2020 - Proceedings of the 2020 ACM Symposium on Cloud Computing (pp. 447–461). Association for Computing Machinery, Inc. https://doi.org/10.1145/3419111.3421296

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free