Unified Kernel-Segregated Transpose Convolution Operation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tida, Vijay Srinivas, Hossen, Md Imran, Shan, Liqun, Chilukoti, Sai Venkatesh, Hsu, Sonya, Hei, Xiali
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913712502210560
author Tida, Vijay Srinivas
Hossen, Md Imran
Shan, Liqun
Chilukoti, Sai Venkatesh
Hsu, Sonya
Hei, Xiali
author_facet Tida, Vijay Srinivas
Hossen, Md Imran
Shan, Liqun
Chilukoti, Sai Venkatesh
Hsu, Sonya
Hei, Xiali
contents The optimization of the transpose convolution layer for deep learning applications is achieved with the kernel segregation mechanism. However, kernel segregation has disadvantages, such as computing extra elements to obtain the output feature map with odd dimensions while launching a thread. To mitigate this problem, we introduce a unified kernel segregation approach that limits the usage of memory and computational resources by employing one unified kernel to execute four sub-kernels. The findings reveal that the suggested approach achieves an average computational speedup of 2.03x (3.89x) when tested on specific datasets with an RTX 2070 GPU (Intel Xeon CPU). The ablation study shows an average computational speedup of 3.5x when evaluating the transpose convolution layers from well-known Generative Adversarial Networks (GANs). The implementation of the proposed method for the transpose convolution layers in the EB-GAN model demonstrates significant memory savings of up to 35 MB.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20493
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unified Kernel-Segregated Transpose Convolution Operation
Tida, Vijay Srinivas
Hossen, Md Imran
Shan, Liqun
Chilukoti, Sai Venkatesh
Hsu, Sonya
Hei, Xiali
Machine Learning
Artificial Intelligence
The optimization of the transpose convolution layer for deep learning applications is achieved with the kernel segregation mechanism. However, kernel segregation has disadvantages, such as computing extra elements to obtain the output feature map with odd dimensions while launching a thread. To mitigate this problem, we introduce a unified kernel segregation approach that limits the usage of memory and computational resources by employing one unified kernel to execute four sub-kernels. The findings reveal that the suggested approach achieves an average computational speedup of 2.03x (3.89x) when tested on specific datasets with an RTX 2070 GPU (Intel Xeon CPU). The ablation study shows an average computational speedup of 3.5x when evaluating the transpose convolution layers from well-known Generative Adversarial Networks (GANs). The implementation of the proposed method for the transpose convolution layers in the EB-GAN model demonstrates significant memory savings of up to 35 MB.
title Unified Kernel-Segregated Transpose Convolution Operation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.20493