Self-supervised pretraining for an iterative image size agnostic vision transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Prisadnikov, Nedyalko, Paudel, Danda Pani, Fu, Yuqian, Van Gool, Luc
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911614224039936
author Prisadnikov, Nedyalko
Paudel, Danda Pani
Fu, Yuqian
Van Gool, Luc
author_facet Prisadnikov, Nedyalko
Paudel, Danda Pani
Fu, Yuqian
Van Gool, Luc
contents Vision Transformers (ViTs) dominate self-supervised learning (SSL). While they have proven highly effective for large-scale pretraining, they are computationally inefficient and scale poorly with image size. Consequently, foundational models like DINO are constrained to low-resolution processing. A recent foveal-inspired transformer achieves resolution agnosticism by iteratively processing a fixed-size context of multi-zoom patches. This model demonstrated promising results via supervised learning, utilizing a sequential, recurrent-like process without backpropagation through time. To unlock its potential as a foundational backbone, we introduce a novel sequential-to-global SSL framework based on DINO's self-distillation objective. Supported by an efficient integral-image patch extraction method, our approach enables large-scale pretraining for image-size agnostic vision encoders. We achieve competitive performance on ImageNet-1K and downstream classification tasks, maintaining a constant computational budget regardless of input resolution.
format Preprint
id arxiv_https___arxiv_org_abs_2604_20392
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Self-supervised pretraining for an iterative image size agnostic vision transformer
Prisadnikov, Nedyalko
Paudel, Danda Pani
Fu, Yuqian
Van Gool, Luc
Computer Vision and Pattern Recognition
Vision Transformers (ViTs) dominate self-supervised learning (SSL). While they have proven highly effective for large-scale pretraining, they are computationally inefficient and scale poorly with image size. Consequently, foundational models like DINO are constrained to low-resolution processing. A recent foveal-inspired transformer achieves resolution agnosticism by iteratively processing a fixed-size context of multi-zoom patches. This model demonstrated promising results via supervised learning, utilizing a sequential, recurrent-like process without backpropagation through time. To unlock its potential as a foundational backbone, we introduce a novel sequential-to-global SSL framework based on DINO's self-distillation objective. Supported by an efficient integral-image patch extraction method, our approach enables large-scale pretraining for image-size agnostic vision encoders. We achieve competitive performance on ImageNet-1K and downstream classification tasks, maintaining a constant computational budget regardless of input resolution.
title Self-supervised pretraining for an iterative image size agnostic vision transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.20392