bViT: Investigating Single-Block Recurrence in Vision Transformers for Image Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Byra, Michal, Olszowiec, Pawel, Stefanski, Grzegorz, Gruszczynski, Grzegorz, Presta, Alberto
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913112744001536
author Byra, Michal
Olszowiec, Pawel
Stefanski, Grzegorz
Gruszczynski, Grzegorz
Presta, Alberto
author_facet Byra, Michal
Olszowiec, Pawel
Stefanski, Grzegorz
Gruszczynski, Grzegorz
Presta, Alberto
contents Vision Transformers (ViTs) are built by stacking independently parameterized blocks, but it remains unclear how much of this depth requires layer specific transformations and how much can be realized through recurrent computation. We study this question with bViT, a single-block recurrent ViT in which one transformer block is applied repeatedly to process an image. This architecture preserves the iterative structure of a deep ViT while removing layer specific block parameterization, providing a controlled setting for studying recurrence in vision. On ImageNet-1K, a 12-step bViT-B achieves accuracy comparable to standard ViT-B under the same training recipe and computational budget, while using an order of magnitude fewer parameters. We observe that recurrent performance improves with representation width, with wider bViTs recovering much more of the performance of standard ViTs than narrow variants. We interpret this behavior as implicit depth multiplexing, where a shared block expresses multiple step-dependent computations through the evolving hidden state. Beyond ImageNet classification, bViT transfers competitively to downstream tasks and enables parameter-efficient fine-tuning. Mechanistic analyses of activations, attention and step-specific pruning show that the shared block changes its effective behavior across recurrent steps rather than simply repeating the same computation. Our results suggest that a large fraction of ViT depth can be implemented through recurrent reuse, provided that the representation space is sufficiently wide.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10661
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle bViT: Investigating Single-Block Recurrence in Vision Transformers for Image Recognition
Byra, Michal
Olszowiec, Pawel
Stefanski, Grzegorz
Gruszczynski, Grzegorz
Presta, Alberto
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision Transformers (ViTs) are built by stacking independently parameterized blocks, but it remains unclear how much of this depth requires layer specific transformations and how much can be realized through recurrent computation. We study this question with bViT, a single-block recurrent ViT in which one transformer block is applied repeatedly to process an image. This architecture preserves the iterative structure of a deep ViT while removing layer specific block parameterization, providing a controlled setting for studying recurrence in vision. On ImageNet-1K, a 12-step bViT-B achieves accuracy comparable to standard ViT-B under the same training recipe and computational budget, while using an order of magnitude fewer parameters. We observe that recurrent performance improves with representation width, with wider bViTs recovering much more of the performance of standard ViTs than narrow variants. We interpret this behavior as implicit depth multiplexing, where a shared block expresses multiple step-dependent computations through the evolving hidden state. Beyond ImageNet classification, bViT transfers competitively to downstream tasks and enables parameter-efficient fine-tuning. Mechanistic analyses of activations, attention and step-specific pruning show that the shared block changes its effective behavior across recurrent steps rather than simply repeating the same computation. Our results suggest that a large fraction of ViT depth can be implemented through recurrent reuse, provided that the representation space is sufficiently wide.
title bViT: Investigating Single-Block Recurrence in Vision Transformers for Image Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.10661