Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wijeratne, Sasindu, Kannan, Rajgopal, Prasanna, Viktor
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908280326979584
author Wijeratne, Sasindu
Kannan, Rajgopal
Prasanna, Viktor
author_facet Wijeratne, Sasindu
Kannan, Rajgopal
Prasanna, Viktor
contents Sparse Matricized Tensor Times Khatri-Rao Product (spMTTKRP) is the bottleneck kernel of sparse tensor decomposition. In tensor decomposition, spMTTKRP is performed iteratively along all the modes of an input tensor. In this work, we propose a mode-specific tensor layout on GPU that uses multiple tensor copies, where each copy is optimized for a specific mode. The proposed tensor layout increases the data locality of external memory accesses and eliminates the intermediate values communicated between the GPU thread blocks and the GPU global memory. We also propose a tensor partitioning scheme to optimally distribute the total computations among GPU streaming multiprocessors based on the sparsity and the dimensions of the input tensor. Our approach achieves a geometric mean speedup of 2.4x, 7.9x, and 8.9x in total execution time compared with the state-of-the-art GPU baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18198
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
Wijeratne, Sasindu
Kannan, Rajgopal
Prasanna, Viktor
Distributed, Parallel, and Cluster Computing
Sparse Matricized Tensor Times Khatri-Rao Product (spMTTKRP) is the bottleneck kernel of sparse tensor decomposition. In tensor decomposition, spMTTKRP is performed iteratively along all the modes of an input tensor. In this work, we propose a mode-specific tensor layout on GPU that uses multiple tensor copies, where each copy is optimized for a specific mode. The proposed tensor layout increases the data locality of external memory accesses and eliminates the intermediate values communicated between the GPU thread blocks and the GPU global memory. We also propose a tensor partitioning scheme to optimally distribute the total computations among GPU streaming multiprocessors based on the sparsity and the dimensions of the input tensor. Our approach achieves a geometric mean speedup of 2.4x, 7.9x, and 8.9x in total execution time compared with the state-of-the-art GPU baselines.
title Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2503.18198