A GPU scheduling framework to accelerate hyper-parameter optimization in deep learning clusters

4Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.

Abstract

This paper proposes Hermes, a container-based preemptive GPU scheduling framework for accelerating hyper-parameter optimization in deep learning (DL) clusters. Hermes accelerates hyper-parameter optimization by time-sharing between DL jobs and prioritizing jobs with more promising hyper-parameter combinations. Hermes’s scheduling policy is grounded on the observation that good hyper-parameter combinations converge quickly in the early phases of training. By giving higher priority to fast-converging containers, Hermes’s GPU preemption mechanism can accelerate training. This enables users to find optimal hyper-parameters faster without losing the progress of a container. We have implemented Hermes over Kubernetes and compared its performance against existing scheduling frameworks. Experiments show that Hermes reduces the time for hyper-parameter optimization up to 4.04 times against previously proposed scheduling policies such as FIFO, round-robin (RR), and SLAQ, with minimal time-sharing overhead.

Cite

CITATION STYLE

APA

Son, J., Yoo, Y., Kim, K. R., Kim, Y., Lee, K., & Park, S. (2021). A GPU scheduling framework to accelerate hyper-parameter optimization in deep learning clusters. Electronics (Switzerland), 10(3), 1–15. https://doi.org/10.3390/electronics10030350

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free