Abstract
Deep learning models having transpose convolution layers requires optimization to deploy in the resource constraint Internet of Things (IoT) devices. The main reason is the presence of zeros at predefined positions in the input feature maps after upsampling layer leads to higher memory and computation load requirements for transpose convolution operations. We propose an algorithmic-level optimization technique based on kernel segregation mechanisms for efficient transpose convolution implementation to address these issues without needing an upsampling layer. Experimental results showed that the proposed approach showed an average of 3.7 × (3.4×) faster computation using an Intel Xeon CPU (RTX 2070 GPU) than the conventional method. Further, we analyzed the performance using different transpose convolution layers from the popular Generative Adversarial Network (GAN) models and a simple deep learning model with one transpose convolution layer. There is a significant improvement in computation speed and substantial memory savings from the obtained results.
Author supplied keywords
Cite
CITATION STYLE
Tida, V. S., Hsu, S., Chilukoti, S. V., & Hei, X. (2023). Kernel-Segregated Transpose Convolution Operation. In Proceedings of the Annual Hawaii International Conference on System Sciences (Vol. 2023-January, pp. 6934–6943). IEEE Computer Society. https://doi.org/10.24251/hicss.2023.840
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.