A Computationally Efficient Multidimensional Vision Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ichi, Alaa El, Jbilou, Khalide
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911463085441024
author Ichi, Alaa El
Jbilou, Khalide
author_facet Ichi, Alaa El
Jbilou, Khalide
contents Vision Transformers have achieved state-of-the-art performance in a wide range of computer vision tasks, but their practical deployment is limited by high computational and memory costs. In this paper, we introduce a novel tensor-based framework for Vision Transformers built upon the Tensor Cosine Product (Cproduct). By exploiting multilinear structures inherent in image data and the orthogonality of cosine transforms, the proposed approach enables efficient attention mechanisms and structured feature representations. We develop the theoretical foundations of the tensor cosine product, analyze its algebraic properties, and integrate it into a new Cproduct-based Vision Transformer architecture (TCP-ViT). Numerical experiments on standard classification and segmentation benchmarks demonstrate that the proposed method achieves a uniform 1/C parameter reduction (where C is the number of channels) while maintaining competitive accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2602_19982
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Computationally Efficient Multidimensional Vision Transformer
Ichi, Alaa El
Jbilou, Khalide
Machine Learning
Numerical Analysis
Vision Transformers have achieved state-of-the-art performance in a wide range of computer vision tasks, but their practical deployment is limited by high computational and memory costs. In this paper, we introduce a novel tensor-based framework for Vision Transformers built upon the Tensor Cosine Product (Cproduct). By exploiting multilinear structures inherent in image data and the orthogonality of cosine transforms, the proposed approach enables efficient attention mechanisms and structured feature representations. We develop the theoretical foundations of the tensor cosine product, analyze its algebraic properties, and integrate it into a new Cproduct-based Vision Transformer architecture (TCP-ViT). Numerical experiments on standard classification and segmentation benchmarks demonstrate that the proposed method achieves a uniform 1/C parameter reduction (where C is the number of channels) while maintaining competitive accuracy.
title A Computationally Efficient Multidimensional Vision Transformer
topic Machine Learning
Numerical Analysis
url https://arxiv.org/abs/2602.19982