ViT Registers and Fractal ViT

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chou, Jason Chuan-Chih, Kumar, Abhinav, Garg, Shivank
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917216814891008
author Chou, Jason Chuan-Chih
Kumar, Abhinav
Garg, Shivank
author_facet Chou, Jason Chuan-Chih
Kumar, Abhinav
Garg, Shivank
contents Drawing inspiration from recent findings including surprisingly decent performance of transformers without positional encoding (NoPE) in the domain of language models and how registers (additional throwaway tokens not tied to input) may improve the performance of large vision transformers (ViTs), we invent and test a variant of ViT called fractal ViT that breaks permutation invariance among the tokens by applying an attention mask between the regular tokens and ``summary tokens'' similar to registers, in isolation or in combination with various positional encodings. These models do not improve upon ViT with registers, highlighting the fact that these findings may be scale, domain, or application-specific.
format Preprint
id arxiv_https___arxiv_org_abs_2601_15506
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ViT Registers and Fractal ViT
Chou, Jason Chuan-Chih
Kumar, Abhinav
Garg, Shivank
Computation and Language
Machine Learning
Drawing inspiration from recent findings including surprisingly decent performance of transformers without positional encoding (NoPE) in the domain of language models and how registers (additional throwaway tokens not tied to input) may improve the performance of large vision transformers (ViTs), we invent and test a variant of ViT called fractal ViT that breaks permutation invariance among the tokens by applying an attention mask between the regular tokens and ``summary tokens'' similar to registers, in isolation or in combination with various positional encodings. These models do not improve upon ViT with registers, highlighting the fact that these findings may be scale, domain, or application-specific.
title ViT Registers and Fractal ViT
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.15506