Abstract
Deep learning (DL) models have rapidly evolved, and their scales have become larger. Pipeline parallelism is used to execute a large-scale DL model. In pipeline parallelism, DL models are partitioned and allocated graphics processing units (GPUs) to execute each partition. However, to execute numerous DL services in clusters with heterogeneous GPUs, a DL model partitioning that considers a specific type and number of GPUs available and the GPUs allocated for the service is required. We propose resource aware model partitioning and allocation (RAMPA) to execute more DL services while satisfying the performance requirements. RAMPA minimizes the allocation of important resources for future DL service execution to avoid inhibiting future DL service execution. We define the resource allocation cost based on resource importance. Furthermore, we formulate the impact of model partitioning and allocated resources on service performance. We define an optimization problem to minimize resource allocation costs while satisfying service performance requirements. We evaluated the effectiveness of RAMPA by simulating the execution of DL services in clusters with heterogeneous GPUs. The results demonstrate that more services can be executed while satisfying performance requirements compared to the conventional method. RAMPA enabled efficient GPU utilization to deliver many DL services.
Author supplied keywords
Cite
CITATION STYLE
Ikoma, A., Ohsita, Y., & Murata, M. (2025). Resource Aware Deep Learning Model Partitioning and Allocation for Inference Task in Clusters With Heterogeneous Graphics Processing Units. IEEE Transactions on Cloud Computing, 13(4), 1105–1118. https://doi.org/10.1109/TCC.2025.3626959
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.