Abstract
This paper proposes Hermes, a container-based preemptive GPU scheduling framework for accelerating hyper-parameter optimization in deep learning (DL) clusters. Hermes accelerates hyper-parameter optimization by time-sharing between DL jobs and prioritizing jobs with more promising hyper-parameter combinations. Hermes’s scheduling policy is grounded on the observation that good hyper-parameter combinations converge quickly in the early phases of training. By giving higher priority to fast-converging containers, Hermes’s GPU preemption mechanism can accelerate training. This enables users to find optimal hyper-parameters faster without losing the progress of a container. We have implemented Hermes over Kubernetes and compared its performance against existing scheduling frameworks. Experiments show that Hermes reduces the time for hyper-parameter optimization up to 4.04 times against previously proposed scheduling policies such as FIFO, round-robin (RR), and SLAQ, with minimal time-sharing overhead.
Author supplied keywords
Cite
CITATION STYLE
Son, J., Yoo, Y., Kim, K. R., Kim, Y., Lee, K., & Park, S. (2021). A GPU scheduling framework to accelerate hyper-parameter optimization in deep learning clusters. Electronics (Switzerland), 10(3), 1–15. https://doi.org/10.3390/electronics10030350
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.