Abstract
Machine Learning workloads such as deep learning and hyperparameter tuning are compute-intensive by nature. Parallel execution is key to reducing the learning time. The Ray Framework is a distributed middleware that provides primitives to seamlessly parallelize machine learning code execution across a cluster of compute node. Launching a Ray managed machine learning application requires a Ray cluster that is diligently configured, well connected and easily scalable. Kubernetes, the container management middleware, satisfies all the requirements to create and scale ray clusters. However, setting up a cluster within Kubernetes, is a tedious and error prone task when done manually. In this paper we present KubeRay, an Operator and suite of tools designed, and built to create Ray cluster in Kubernetes with minimum effort. We present our architectural choices, our open-source implementation, and we analyze the performance of our solution.
Author supplied keywords
Cite
CITATION STYLE
Kanso, A., Palencia, E., Patra, K., Shan, J., Chao, M., Wei, X., … Qiao, S. (2021). Designing a kubernetes operator for machine learning applications. In WoC 2021 - Proceedings of the 2021 7th International Workshop on Container Technologies and Container Clouds (pp. 7–12). Association for Computing Machinery, Inc. https://doi.org/10.1145/3493649.3493654
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.