Image Recognition with Online Lightweight Vision Transformer: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zherui, Xu, Rongtao, Zhou, Jie, Wang, Changwei, Pei, Xingtian, Xu, Wenhao, Zhang, Jiguang, Guo, Li, Gao, Longxiang, Xu, Wenbo, Xu, Shibiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915515610431488
author Zhang, Zherui
Xu, Rongtao
Zhou, Jie
Wang, Changwei
Pei, Xingtian
Xu, Wenhao
Zhang, Jiguang
Guo, Li
Gao, Longxiang
Xu, Wenbo
Xu, Shibiao
author_facet Zhang, Zherui
Xu, Rongtao
Zhou, Jie
Wang, Changwei
Pei, Xingtian
Xu, Wenhao
Zhang, Jiguang
Guo, Li
Gao, Longxiang
Xu, Wenbo
Xu, Shibiao
contents The Transformer architecture has achieved significant success in natural language processing, motivating its adaptation to computer vision tasks. Unlike convolutional neural networks, vision transformers inherently capture long-range dependencies and enable parallel processing, yet lack inductive biases and efficiency benefits, facing significant computational and memory challenges that limit its real-world applicability. This paper surveys various online strategies for generating lightweight vision transformers for image recognition, focusing on three key areas: Efficient Component Design, Dynamic Network, and Knowledge Distillation. We evaluate the relevant exploration for each topic on the ImageNet-1K benchmark, analyzing trade-offs among precision, parameters, throughput, and more to highlight their respective advantages, disadvantages, and flexibility. Finally, we propose future research directions and potential challenges in the lightweighting of vision transformers with the aim of inspiring further exploration and providing practical guidance for the community. Project Page: https://github.com/ajxklo/Lightweight-VIT
format Preprint
id arxiv_https___arxiv_org_abs_2505_03113
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Image Recognition with Online Lightweight Vision Transformer: A Survey
Zhang, Zherui
Xu, Rongtao
Zhou, Jie
Wang, Changwei
Pei, Xingtian
Xu, Wenhao
Zhang, Jiguang
Guo, Li
Gao, Longxiang
Xu, Wenbo
Xu, Shibiao
Computer Vision and Pattern Recognition
The Transformer architecture has achieved significant success in natural language processing, motivating its adaptation to computer vision tasks. Unlike convolutional neural networks, vision transformers inherently capture long-range dependencies and enable parallel processing, yet lack inductive biases and efficiency benefits, facing significant computational and memory challenges that limit its real-world applicability. This paper surveys various online strategies for generating lightweight vision transformers for image recognition, focusing on three key areas: Efficient Component Design, Dynamic Network, and Knowledge Distillation. We evaluate the relevant exploration for each topic on the ImageNet-1K benchmark, analyzing trade-offs among precision, parameters, throughput, and more to highlight their respective advantages, disadvantages, and flexibility. Finally, we propose future research directions and potential challenges in the lightweighting of vision transformers with the aim of inspiring further exploration and providing practical guidance for the community. Project Page: https://github.com/ajxklo/Lightweight-VIT
title Image Recognition with Online Lightweight Vision Transformer: A Survey
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.03113