DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Foundation Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yin, Shicheng, Yin, Kaixuan, Liu, Yang, Chen, Weixing, Lin, Liang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914063664021504
author Yin, Shicheng
Yin, Kaixuan
Liu, Yang
Chen, Weixing
Lin, Liang
author_facet Yin, Shicheng
Yin, Kaixuan
Liu, Yang
Chen, Weixing
Lin, Liang
contents The content-agnostic, fixed-grid tokenizers used by standard large-scale vision models like Vision Transformer (ViT) and Vision Mamba (Vim) represent a fundamental performance bottleneck, creating a trade-off between capturing fine-grained detail and suffering from redundant computation. To resolve this dilemma, we introduce DART, a fully differentiable Dynamic Adaptive Region Tokenizer. DART employs learnable region scores and quantile-based partitioning to create content-aware patches of varying sizes, intelligently allocating a higher token density to information-rich regions. The impact of this approach is profound: it unlocks a more intelligent scaling paradigm, where a DART-equipped DeiT-Small (22M parameters) matches the performance of a DeiT-Base (86M) with nearly double the inference speed by efficiently capturing high-resolution details in key regions. Furthermore, the principle of adaptive tokenization proves its generality with clear benefits in dense prediction and spatiotemporal video tasks. We argue that by resolving the tokenizer bottleneck at its source, adaptive tokenization is a key component for building the next generation of more efficient and capable foundation models for multimodal AI, robotics, and content generation. Code is available at https://github.com/HCPLab-SYSU/DART.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10390
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Foundation Models
Yin, Shicheng
Yin, Kaixuan
Liu, Yang
Chen, Weixing
Lin, Liang
Computer Vision and Pattern Recognition
The content-agnostic, fixed-grid tokenizers used by standard large-scale vision models like Vision Transformer (ViT) and Vision Mamba (Vim) represent a fundamental performance bottleneck, creating a trade-off between capturing fine-grained detail and suffering from redundant computation. To resolve this dilemma, we introduce DART, a fully differentiable Dynamic Adaptive Region Tokenizer. DART employs learnable region scores and quantile-based partitioning to create content-aware patches of varying sizes, intelligently allocating a higher token density to information-rich regions. The impact of this approach is profound: it unlocks a more intelligent scaling paradigm, where a DART-equipped DeiT-Small (22M parameters) matches the performance of a DeiT-Base (86M) with nearly double the inference speed by efficiently capturing high-resolution details in key regions. Furthermore, the principle of adaptive tokenization proves its generality with clear benefits in dense prediction and spatiotemporal video tasks. We argue that by resolving the tokenizer bottleneck at its source, adaptive tokenization is a key component for building the next generation of more efficient and capable foundation models for multimodal AI, robotics, and content generation. Code is available at https://github.com/HCPLab-SYSU/DART.
title DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Foundation Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.10390