Abstract
Computations performed by using convolutional layers in deep learning require significant resources; thus, their scope of applicability is limited. When deep neural network models are employed in an edge-computing system, the limited computational power and storage resources of edge devices can degrade inference performance, require a considerable amount of computation time, and result in increased energy consumption. To address these issues, this study presents a convolutional-layer partitioning model, based on the fused tile partitioning (FTP) algorithm, for enhancing the distributed inference capabilities of edge devices. First, a resource-adaptive workload-partitioning optimization model is designed to promote load balancing across heterogeneous edge systems. Next, the FTP algorithm is improved, leading to a new layer-fused partitioning method that is used to solve the optimization model. The results of simulation experiments show that the proposed convolutional-layer partitioning method effectively improves the inference performance of edge devices. When five edge devices are used, the speed of the proposed method becomes 1.65–3.48 times those of existing algorithms.
Author supplied keywords
Cite
CITATION STYLE
Yuan, Q., & Li, Z. (2025). Distributed Inference Models and Algorithms for Heterogeneous Edge Systems Using Deep Learning. Applied Sciences (Switzerland), 15(3). https://doi.org/10.3390/app15031097
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.