Kernel-Segregated Transpose Convolution Operation

0Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.

Abstract

Deep learning models having transpose convolution layers requires optimization to deploy in the resource constraint Internet of Things (IoT) devices. The main reason is the presence of zeros at predefined positions in the input feature maps after upsampling layer leads to higher memory and computation load requirements for transpose convolution operations. We propose an algorithmic-level optimization technique based on kernel segregation mechanisms for efficient transpose convolution implementation to address these issues without needing an upsampling layer. Experimental results showed that the proposed approach showed an average of 3.7 × (3.4×) faster computation using an Intel Xeon CPU (RTX 2070 GPU) than the conventional method. Further, we analyzed the performance using different transpose convolution layers from the popular Generative Adversarial Network (GAN) models and a simple deep learning model with one transpose convolution layer. There is a significant improvement in computation speed and substantial memory savings from the obtained results.

Cite

CITATION STYLE

APA

Tida, V. S., Hsu, S., Chilukoti, S. V., & Hei, X. (2023). Kernel-Segregated Transpose Convolution Operation. In Proceedings of the Annual Hawaii International Conference on System Sciences (Vol. 2023-January, pp. 6934–6943). IEEE Computer Society. https://doi.org/10.24251/hicss.2023.840

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free