JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zou, Dongyun, Zhang, Zhuoyang, Chen, Junyu, He, Wenkun, Peng, Qinhe, Ye, Hanrong, Lu, Yao, Yin, Hongxu, Wang, Yu, Han, Song, Cai, Han
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910258244354048
author Zou, Dongyun
Zhang, Zhuoyang
Chen, Junyu
He, Wenkun
Peng, Qinhe
Ye, Hanrong
Lu, Yao
Yin, Hongxu
Wang, Yu
Han, Song
Cai, Han
author_facet Zou, Dongyun
Zhang, Zhuoyang
Chen, Junyu
He, Wenkun
Peng, Qinhe
Ye, Hanrong
Lu, Yao
Yin, Hongxu
Wang, Yu
Han, Song
Cai, Han
contents We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention vision foundation models while achieving substantially higher inference efficiency on high-resolution images. At the core of our approach is Post-Training Attention Search, a post-training acceleration framework that converts pre-trained full-attention ViTs into efficient hybrid-attention variants by identifying and replacing redundant full-attention blocks with linear or window-attention blocks. By inheriting the MLP and attention weights from the base model, Post-Training Attention Search efficiently explores the architectural design space through three key steps: (1) optimizing the linear-attention block design; (2) finding the best combination of linear-attention and window-attention blocks; and (3) identifying and preserving critical full-attention blocks. We evaluate JetViT on two representative high-resolution vision foundation models, DINOv3 and DepthAnythingV2. On the NVIDIA H100 GPU, JetViT achieves up to 1.79x higher throughput and up to 44.81% lower latency without sacrificing accuracy. We will release our code and accelerated ViT models soon.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26636
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
Zou, Dongyun
Zhang, Zhuoyang
Chen, Junyu
He, Wenkun
Peng, Qinhe
Ye, Hanrong
Lu, Yao
Yin, Hongxu
Wang, Yu
Han, Song
Cai, Han
Computer Vision and Pattern Recognition
Artificial Intelligence
We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention vision foundation models while achieving substantially higher inference efficiency on high-resolution images. At the core of our approach is Post-Training Attention Search, a post-training acceleration framework that converts pre-trained full-attention ViTs into efficient hybrid-attention variants by identifying and replacing redundant full-attention blocks with linear or window-attention blocks. By inheriting the MLP and attention weights from the base model, Post-Training Attention Search efficiently explores the architectural design space through three key steps: (1) optimizing the linear-attention block design; (2) finding the best combination of linear-attention and window-attention blocks; and (3) identifying and preserving critical full-attention blocks. We evaluate JetViT on two representative high-resolution vision foundation models, DINOv3 and DepthAnythingV2. On the NVIDIA H100 GPU, JetViT achieves up to 1.79x higher throughput and up to 44.81% lower latency without sacrificing accuracy. We will release our code and accelerated ViT models soon.
title JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.26636