ViT-5: Vision Transformers for The Mid-2020s
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Feng, Ren, Sucheng, Zhang, Tiezheng, Neskovic, Predrag, Bhattad, Anand, Xie, Cihang, Yuille, Alan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
by: Yang, Timing, et al.
Published: (2025)
by: Yang, Timing, et al.
Published: (2025)
M-VAR: Decoupled Scale-wise Autoregressive Modeling for High-Quality Image Generation
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
Rejuvenating image-GPT as Strong Visual Representation Learners
by: Ren, Sucheng, et al.
Published: (2023)
by: Ren, Sucheng, et al.
Published: (2023)
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
Mamba-R: Vision Mamba ALSO Needs Registers
by: Wang, Feng, et al.
Published: (2024)
by: Wang, Feng, et al.
Published: (2024)
SPFormer: Enhancing Vision Transformer with Superpixel Representation
by: Mei, Jieru, et al.
Published: (2024)
by: Mei, Jieru, et al.
Published: (2024)
Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
by: Liu, Haoyu, et al.
Published: (2026)
by: Liu, Haoyu, et al.
Published: (2026)
Adventurer: Optimizing Vision Mamba Architecture Designs for Efficiency
by: Wang, Feng, et al.
Published: (2024)
by: Wang, Feng, et al.
Published: (2024)
Autoregressive Pretraining with Mamba in Vision
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
Name That Part: 3D Part Segmentation and Naming
by: Paul, Soumava, et al.
Published: (2025)
by: Paul, Soumava, et al.
Published: (2025)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
by: Zhu, Chen, et al.
Published: (2025)
by: Zhu, Chen, et al.
Published: (2025)
From Pixels to Objects: A Hierarchical Approach for Part and Object Segmentation Using Local and Global Aggregation
by: Xie, Yunfei, et al.
Published: (2024)
by: Xie, Yunfei, et al.
Published: (2024)
ACC-ViT : Atrous Convolution's Comeback in Vision Transformers
by: Ibtehaz, Nabil, et al.
Published: (2024)
by: Ibtehaz, Nabil, et al.
Published: (2024)
ADFQ-ViT: Activation-Distribution-Friendly Post-Training Quantization for Vision Transformers
by: Jiang, Yanfeng, et al.
Published: (2024)
by: Jiang, Yanfeng, et al.
Published: (2024)
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
by: Ma, Xiaochen, et al.
Published: (2023)
by: Ma, Xiaochen, et al.
Published: (2023)
Dictionary-based Framework for Interpretable and Consistent Object Parsing
by: Zhang, Tiezheng, et al.
Published: (2025)
by: Zhang, Tiezheng, et al.
Published: (2025)
ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions
by: Xia, Chunlong, et al.
Published: (2024)
by: Xia, Chunlong, et al.
Published: (2024)
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
ViT-Explainer: An Interactive Walkthrough of the Vision Transformer Pipeline
by: Hernandez, Juan Manuel, et al.
Published: (2026)
by: Hernandez, Juan Manuel, et al.
Published: (2026)
Filter & Align: Leveraging Human Knowledge to Curate Image-Text Data
by: Zhang, Lei, et al.
Published: (2023)
by: Zhang, Lei, et al.
Published: (2023)
ViT-FIQA: Assessing Face Image Quality using Vision Transformers
by: Atzori, Andrea, et al.
Published: (2025)
by: Atzori, Andrea, et al.
Published: (2025)
MPTQ-ViT: Mixed-Precision Post-Training Quantization for Vision Transformer
by: Tai, Yu-Shan, et al.
Published: (2024)
by: Tai, Yu-Shan, et al.
Published: (2024)
ViT-1.58b: Mobile Vision Transformers in the 1-bit Era
by: Yuan, Zhengqing, et al.
Published: (2024)
by: Yuan, Zhengqing, et al.
Published: (2024)
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs
by: Yao, Ting, et al.
Published: (2024)
by: Yao, Ting, et al.
Published: (2024)
VAT: Vision Action Transformer by Unlocking Full Representation of ViT
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
by: Li, Zhuowan, et al.
Published: (2022)
by: Li, Zhuowan, et al.
Published: (2022)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
by: Chen, Jieneng, et al.
Published: (2024)
by: Chen, Jieneng, et al.
Published: (2024)
ViT-DD: Multi-Task Vision Transformer for Semi-Supervised Driver Distraction Detection
by: Ma, Yunsheng, et al.
Published: (2022)
by: Ma, Yunsheng, et al.
Published: (2022)
APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision Transformers
by: Wu, Zhuguanyu, et al.
Published: (2025)
by: Wu, Zhuguanyu, et al.
Published: (2025)
CAS-ViT: Convolutional Additive Self-attention Vision Transformers for Efficient Mobile Applications
by: Zhang, Tianfang, et al.
Published: (2024)
by: Zhang, Tianfang, et al.
Published: (2024)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
by: Kuzucu, Selim, et al.
Published: (2025)
by: Kuzucu, Selim, et al.
Published: (2025)
Hyb-KAN ViT: Hybrid Kolmogorov-Arnold Networks Augmented Vision Transformer
by: Dey, Sainath, et al.
Published: (2025)
by: Dey, Sainath, et al.
Published: (2025)
SVD-ViT: Does SVD Make Vision Transformers Attend More to the Foreground?
by: Murata, Haruhiko, et al.
Published: (2026)
by: Murata, Haruhiko, et al.
Published: (2026)
ViT$^3$: Unlocking Test-Time Training in Vision
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
HResFormer: Hybrid Residual Transformer for Volumetric Medical Image Segmentation
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
ViLBench: A Suite for Vision-Language Process Reward Modeling
by: Tu, Haoqin, et al.
Published: (2025)
by: Tu, Haoqin, et al.
Published: (2025)
LL-ViT: Edge Deployable Vision Transformers with Look Up Table Neurons
by: Nag, Shashank, et al.
Published: (2025)
by: Nag, Shashank, et al.
Published: (2025)
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
by: Zhang, Tiezheng, et al.
Published: (2025)
by: Zhang, Tiezheng, et al.
Published: (2025)
Similar Items
-
ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
by: Yang, Timing, et al.
Published: (2025) -
M-VAR: Decoupled Scale-wise Autoregressive Modeling for High-Quality Image Generation
by: Ren, Sucheng, et al.
Published: (2024) -
Rejuvenating image-GPT as Strong Visual Representation Learners
by: Ren, Sucheng, et al.
Published: (2023) -
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
by: Ren, Sucheng, et al.
Published: (2024) -
Mamba-R: Vision Mamba ALSO Needs Registers
by: Wang, Feng, et al.
Published: (2024)