Alias-Free ViT: Fractional Shift Invariance via Linear Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Michaeli, Hagay, Soudry, Daniel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914116083384320
author Michaeli, Hagay
Soudry, Daniel
author_facet Michaeli, Hagay
Soudry, Daniel
contents Transformers have emerged as a competitive alternative to convnets in vision tasks, yet they lack the architectural inductive bias of convnets, which may hinder their potential performance. Specifically, Vision Transformers (ViTs) are not translation-invariant and are more sensitive to minor image translations than standard convnets. Previous studies have shown, however, that convnets are also not perfectly shift-invariant, due to aliasing in downsampling and nonlinear layers. Consequently, anti-aliasing approaches have been proposed to certify convnets' translation robustness. Building on this line of work, we propose an Alias-Free ViT, which combines two main components. First, it uses alias-free downsampling and nonlinearities. Second, it uses linear cross-covariance attention that is shift-equivariant to both integer and fractional translations, enabling a shift-invariant global representation. Our model maintains competitive performance in image classification and outperforms similar-sized models in terms of robustness to adversarial translations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22673
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Alias-Free ViT: Fractional Shift Invariance via Linear Attention
Michaeli, Hagay
Soudry, Daniel
Computer Vision and Pattern Recognition
Transformers have emerged as a competitive alternative to convnets in vision tasks, yet they lack the architectural inductive bias of convnets, which may hinder their potential performance. Specifically, Vision Transformers (ViTs) are not translation-invariant and are more sensitive to minor image translations than standard convnets. Previous studies have shown, however, that convnets are also not perfectly shift-invariant, due to aliasing in downsampling and nonlinear layers. Consequently, anti-aliasing approaches have been proposed to certify convnets' translation robustness. Building on this line of work, we propose an Alias-Free ViT, which combines two main components. First, it uses alias-free downsampling and nonlinearities. Second, it uses linear cross-covariance attention that is shift-equivariant to both integer and fractional translations, enabling a shift-invariant global representation. Our model maintains competitive performance in image classification and outperforms similar-sized models in terms of robustness to adversarial translations.
title Alias-Free ViT: Fractional Shift Invariance via Linear Attention
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.22673