SPoT: Subpixel Placement of Tokens in Vision Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hjelkrem-Tan, Martine, Aasan, Marius, Arteaga, Gabriel Y., Rivera, Adín Ramírez
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908869421170688
author Hjelkrem-Tan, Martine
Aasan, Marius
Arteaga, Gabriel Y.
Rivera, Adín Ramírez
author_facet Hjelkrem-Tan, Martine
Aasan, Marius
Arteaga, Gabriel Y.
Rivera, Adín Ramírez
contents Vision Transformers naturally accommodate sparsity, yet standard tokenization methods confine features to discrete patch grids. This constraint prevents models from fully exploiting sparse regimes, forcing awkward compromises. We propose Subpixel Placement of Tokens (SPoT), a novel tokenization strategy that positions tokens continuously within images, effectively sidestepping grid-based limitations. With our proposed oracle-guided search, we uncover substantial performance gains achievable with ideal subpixel token positioning, drastically reducing the number of tokens necessary for accurate predictions during inference. SPoT provides a new direction for flexible, efficient, and interpretable ViT architectures, redefining sparsity as a strategic advantage rather than an imposed limitation.
format Preprint
id arxiv_https___arxiv_org_abs_2507_01654
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SPoT: Subpixel Placement of Tokens in Vision Transformers
Hjelkrem-Tan, Martine
Aasan, Marius
Arteaga, Gabriel Y.
Rivera, Adín Ramírez
Computer Vision and Pattern Recognition
Machine Learning
Vision Transformers naturally accommodate sparsity, yet standard tokenization methods confine features to discrete patch grids. This constraint prevents models from fully exploiting sparse regimes, forcing awkward compromises. We propose Subpixel Placement of Tokens (SPoT), a novel tokenization strategy that positions tokens continuously within images, effectively sidestepping grid-based limitations. With our proposed oracle-guided search, we uncover substantial performance gains achievable with ideal subpixel token positioning, drastically reducing the number of tokens necessary for accurate predictions during inference. SPoT provides a new direction for flexible, efficient, and interpretable ViT architectures, redefining sparsity as a strategic advantage rather than an imposed limitation.
title SPoT: Subpixel Placement of Tokens in Vision Transformers
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2507.01654