Designing a kubernetes operator for machine learning applications

12Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Machine Learning workloads such as deep learning and hyperparameter tuning are compute-intensive by nature. Parallel execution is key to reducing the learning time. The Ray Framework is a distributed middleware that provides primitives to seamlessly parallelize machine learning code execution across a cluster of compute node. Launching a Ray managed machine learning application requires a Ray cluster that is diligently configured, well connected and easily scalable. Kubernetes, the container management middleware, satisfies all the requirements to create and scale ray clusters. However, setting up a cluster within Kubernetes, is a tedious and error prone task when done manually. In this paper we present KubeRay, an Operator and suite of tools designed, and built to create Ray cluster in Kubernetes with minimum effort. We present our architectural choices, our open-source implementation, and we analyze the performance of our solution.

Cite

CITATION STYLE

APA

Kanso, A., Palencia, E., Patra, K., Shan, J., Chao, M., Wei, X., … Qiao, S. (2021). Designing a kubernetes operator for machine learning applications. In WoC 2021 - Proceedings of the 2021 7th International Workshop on Container Technologies and Container Clouds (pp. 7–12). Association for Computing Machinery, Inc. https://doi.org/10.1145/3493649.3493654

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free