Abstract
Scheduling computations in each layer of a convolutional neural network on a deep learning (DL) accelerator involves a large number of choices, each of which involves a different set of memory reuse and memory access patterns. Since memory transactions are the primary bottleneck in DL acceleration, these choices can strongly impact the energy and throughput of the accelerator. This work proposes an optimization framework, DeepOpt, for general ASIC-based systolic hardware accelerators for layer-specific and hardware-specific scheduling strategy for each layer of a CNN to optimize energy and latency. Optimal hardware allocation significantly reduces execution cost as compared to generic static hardware resource allocation, e.g., improvements of up to 50 in the energy-delay product for VGG-16 and 41 for GoogleNet-v1.
Author supplied keywords
Cite
CITATION STYLE
Manasi, S. D., & Sapatnekar, S. S. (2021). DeepOpt: Optimized Scheduling of CNNWorkloads for ASIC-based Systolic Deep Learning Accelerators. In Proceedings of the Asia and South Pacific Design Automation Conference, ASP-DAC (pp. 235–241). Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1145/3394885.3431539
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.