A container-based workflow for distributed training of deep learning algorithms in HPC clusters

Jose González-Abad; Álvaro López García; Valentin Y. Kozlov

Journal ArticleOPEN ACCESS

A container-based workflow for distributed training of deep learning algorithms in HPC clusters

Cluster Computing (2023) 26(5) 2815-2834

DOI: 10.1007/s10586-022-03798-7

4Citations

18Readers

Abstract

Deep learning has been postulated as a solution for numerous problems in different branches of science. Given the resource-intensive nature of these models, they often need to be executed on specialized hardware such graphical processing units (GPUs) in a distributed manner. In the academic field, researchers get access to this kind of resources through High Performance Computing (HPC) clusters. This kind of infrastructures make the training of these models difficult due to their multi-user nature and limited user permission. In addition, different HPC clusters may possess different peculiarities that can entangle the research cycle (e.g., libraries dependencies). In this paper we develop a workflow and methodology for the distributed training of deep learning models in HPC clusters which provides researchers with a series of novel advantages. It relies on udocker as containerization tool and on Horovod as library for the distribution of the models across multiple GPUs. udocker does not need any special permission, allowing researchers to run the entire workflow without relying on any administrator. Horovod ensures the efficient distribution of the training independently of the deep learning framework used. Additionally, due to containerization and specific features of the workflow, it provides researchers with a cluster-agnostic way of running their models. The experiments carried out show that the workflow offers good scalability in the distributed training of the models and that it easily adapts to different clusters.

Author supplied keywords

Cite

CITATION STYLE

APA

González-Abad, J., López García, Á., & Kozlov, V. Y. (2023). A container-based workflow for distributed training of deep learning algorithms in HPC clusters. Cluster Computing, 26(5), 2815–2834. https://doi.org/10.1007/s10586-022-03798-7

A container-based workflow for distributed training of deep learning algorithms in HPC clusters

Abstract

Author supplied keywords

Cite

Register to see more suggestions