Octic Vision Transformers: Quicker ViTs Through Equivariance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nordström, David, Edstedt, Johan, Kahl, Fredrik, Bökman, Georg
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911184750379008
author Nordström, David
Edstedt, Johan
Kahl, Fredrik
Bökman, Georg
author_facet Nordström, David
Edstedt, Johan
Kahl, Fredrik
Bökman, Georg
contents Why are state-of-the-art Vision Transformers (ViTs) not designed to exploit natural geometric symmetries such as 90-degree rotations and reflections? In this paper, we argue that there is no fundamental reason, and what has been missing is an efficient implementation. To this end, we introduce Octic Vision Transformers (octic ViTs) which rely on octic group equivariance to capture these symmetries. In contrast to prior equivariant models that increase computational cost, our octic linear layers achieve 5.33x reductions in FLOPs and up to 8x reductions in memory compared to ordinary linear layers. In full octic ViT blocks the computational reductions approach the reductions in the linear layers with increased embedding dimension. We study two new families of ViTs, built from octic blocks, that are either fully octic equivariant or break equivariance in the last part of the network. Training octic ViTs supervised (DeiT-III) and unsupervised (DINOv2) on ImageNet-1K, we find that they match baseline accuracy while at the same time providing substantial efficiency gains.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15441
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Octic Vision Transformers: Quicker ViTs Through Equivariance
Nordström, David
Edstedt, Johan
Kahl, Fredrik
Bökman, Georg
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Why are state-of-the-art Vision Transformers (ViTs) not designed to exploit natural geometric symmetries such as 90-degree rotations and reflections? In this paper, we argue that there is no fundamental reason, and what has been missing is an efficient implementation. To this end, we introduce Octic Vision Transformers (octic ViTs) which rely on octic group equivariance to capture these symmetries. In contrast to prior equivariant models that increase computational cost, our octic linear layers achieve 5.33x reductions in FLOPs and up to 8x reductions in memory compared to ordinary linear layers. In full octic ViT blocks the computational reductions approach the reductions in the linear layers with increased embedding dimension. We study two new families of ViTs, built from octic blocks, that are either fully octic equivariant or break equivariance in the last part of the network. Training octic ViTs supervised (DeiT-III) and unsupervised (DINOv2) on ImageNet-1K, we find that they match baseline accuracy while at the same time providing substantial efficiency gains.
title Octic Vision Transformers: Quicker ViTs Through Equivariance
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.15441