TinyDrop: Tiny Model Guided Token Dropping for Vision Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Guoxin, Wang, Qingyuan, Huang, Binhua, Chen, Shaowu, John, Deepu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914019542040576
author Wang, Guoxin
Wang, Qingyuan
Huang, Binhua
Chen, Shaowu
John, Deepu
author_facet Wang, Guoxin
Wang, Qingyuan
Huang, Binhua
Chen, Shaowu
John, Deepu
contents Vision Transformers (ViTs) achieve strong performance in image classification but incur high computational costs from processing all image tokens. To reduce inference costs in large ViTs without compromising accuracy, we propose TinyDrop, a training-free token dropping framework guided by a lightweight vision model. The guidance model estimates the importance of tokens while performing inference, thereby selectively discarding low-importance tokens if large vit models need to perform attention calculations. The framework operates plug-and-play, requires no architectural modifications, and is compatible with diverse ViT architectures. Evaluations on standard image classification benchmarks demonstrate that our framework reduces FLOPs by up to 80% for ViTs with minimal accuracy degradation, highlighting its generalization capability and practical utility for efficient ViT-based classification.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03379
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TinyDrop: Tiny Model Guided Token Dropping for Vision Transformers
Wang, Guoxin
Wang, Qingyuan
Huang, Binhua
Chen, Shaowu
John, Deepu
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision Transformers (ViTs) achieve strong performance in image classification but incur high computational costs from processing all image tokens. To reduce inference costs in large ViTs without compromising accuracy, we propose TinyDrop, a training-free token dropping framework guided by a lightweight vision model. The guidance model estimates the importance of tokens while performing inference, thereby selectively discarding low-importance tokens if large vit models need to perform attention calculations. The framework operates plug-and-play, requires no architectural modifications, and is compatible with diverse ViT architectures. Evaluations on standard image classification benchmarks demonstrate that our framework reduces FLOPs by up to 80% for ViTs with minimal accuracy degradation, highlighting its generalization capability and practical utility for efficient ViT-based classification.
title TinyDrop: Tiny Model Guided Token Dropping for Vision Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.03379