GenTT: Generate Vectorized Codes for General Tensor Permutation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Yaojian, Ma, Tianyu, Yang, An, Gan, Lin, Zhao, Wenlai, Yang, Guangwen
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916778154655744
author Chen, Yaojian
Ma, Tianyu
Yang, An
Gan, Lin
Zhao, Wenlai
Yang, Guangwen
author_facet Chen, Yaojian
Ma, Tianyu
Yang, An
Gan, Lin
Zhao, Wenlai
Yang, Guangwen
contents Tensor permutation is a fundamental operation widely applied in AI, tensor networks, and related fields. However, it is extremely complex, and different shapes and permutation maps can make a huge difference. SIMD permutation began to be studied in 2006, but the best method at that time was to split complex permutations into multiple simple permutations to do SIMD, which might increase the complexity for very complex permutations. Subsequently, as tensor contraction gained significant attention, researchers explored structured permutations associated with tensor contraction. Progress on general permutations has been limited, and with increasing SIMD bit widths, achieving efficient performance for these permutations has become increasingly challenging. We propose a SIMD permutation toolkit, \system, that generates optimized permutation code for arbitrary instruction sets, bit widths, tensor shapes, and permutation patterns, while maintaining low complexity. In our experiments, \system is able to achieve up to $38\times$ speedup for special cases and $5\times$ for general gases compared to Numpy.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03686
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GenTT: Generate Vectorized Codes for General Tensor Permutation
Chen, Yaojian
Ma, Tianyu
Yang, An
Gan, Lin
Zhao, Wenlai
Yang, Guangwen
Data Structures and Algorithms
Distributed, Parallel, and Cluster Computing
Discrete Mathematics
I.2.2
Tensor permutation is a fundamental operation widely applied in AI, tensor networks, and related fields. However, it is extremely complex, and different shapes and permutation maps can make a huge difference. SIMD permutation began to be studied in 2006, but the best method at that time was to split complex permutations into multiple simple permutations to do SIMD, which might increase the complexity for very complex permutations. Subsequently, as tensor contraction gained significant attention, researchers explored structured permutations associated with tensor contraction. Progress on general permutations has been limited, and with increasing SIMD bit widths, achieving efficient performance for these permutations has become increasingly challenging. We propose a SIMD permutation toolkit, \system, that generates optimized permutation code for arbitrary instruction sets, bit widths, tensor shapes, and permutation patterns, while maintaining low complexity. In our experiments, \system is able to achieve up to $38\times$ speedup for special cases and $5\times$ for general gases compared to Numpy.
title GenTT: Generate Vectorized Codes for General Tensor Permutation
topic Data Structures and Algorithms
Distributed, Parallel, and Cluster Computing
Discrete Mathematics
I.2.2
url https://arxiv.org/abs/2506.03686