Saved in:
| Main Authors: | Chen, Yinjie, Yan, Zipeng, Zhou, Chong, Dai, Bo, Luo, Andrew F. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.21501 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EdgeSAM: Prompt-In-the-Loop Distillation for SAM
by: Zhou, Chong, et al.
Published: (2023)
by: Zhou, Chong, et al.
Published: (2023)
Vision Transformers Need Registers
by: Darcet, Timothée, et al.
Published: (2023)
by: Darcet, Timothée, et al.
Published: (2023)
Vision Transformers Need More Than Registers
by: Shi, Cheng, et al.
Published: (2026)
by: Shi, Cheng, et al.
Published: (2026)
Elastic Attention Cores for Scalable Vision Transformers
by: Song, Alan Z., et al.
Published: (2026)
by: Song, Alan Z., et al.
Published: (2026)
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
by: Zhang, Tianxiao, et al.
Published: (2024)
by: Zhang, Tianxiao, et al.
Published: (2024)
Video Set Distillation: Information Diversification and Temporal Densification
by: Zhao, Yinjie, et al.
Published: (2024)
by: Zhao, Yinjie, et al.
Published: (2024)
Vision Transformers Don't Need Trained Registers
by: Jiang, Nick, et al.
Published: (2025)
by: Jiang, Nick, et al.
Published: (2025)
Leveraging Registers in Vision Transformers for Robust Adaptation
by: Yellapragada, Srikar, et al.
Published: (2025)
by: Yellapragada, Srikar, et al.
Published: (2025)
FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
by: Wang, Zipeng, et al.
Published: (2025)
by: Wang, Zipeng, et al.
Published: (2025)
Mamba-R: Vision Mamba ALSO Needs Registers
by: Wang, Feng, et al.
Published: (2024)
by: Wang, Feng, et al.
Published: (2024)
Generalizing Vision-Language Models with Dedicated Prompt Guidance
by: Li, Xinyao, et al.
Published: (2025)
by: Li, Xinyao, et al.
Published: (2025)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
by: Yoon, Hee Suk, et al.
Published: (2026)
by: Yoon, Hee Suk, et al.
Published: (2026)
Dynamic-Aware Video Distillation: Optimizing Temporal Resolution Based on Video Semantics
by: Zhao, Yinjie, et al.
Published: (2025)
by: Zhao, Yinjie, et al.
Published: (2025)
Distill-SODA: Distilling Self-Supervised Vision Transformer for Source-Free Open-Set Domain Adaptation in Computational Pathology
by: Vray, Guillaume, et al.
Published: (2023)
by: Vray, Guillaume, et al.
Published: (2023)
Zero-Ablation Overstates Register Content Dependence in DINO Vision Transformers
by: Parodi, Felipe, et al.
Published: (2026)
by: Parodi, Felipe, et al.
Published: (2026)
Distillation Dynamics: Towards Understanding Feature-Based Distillation in Vision Transformers
by: Tian, Huiyuan, et al.
Published: (2025)
by: Tian, Huiyuan, et al.
Published: (2025)
Registers Matter for Pixel-Space Diffusion Transformers
by: Starodubcev, Nikita, et al.
Published: (2026)
by: Starodubcev, Nikita, et al.
Published: (2026)
Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression
by: Ye, Hancheng, et al.
Published: (2024)
by: Ye, Hancheng, et al.
Published: (2024)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
by: Zou, Bo, et al.
Published: (2024)
by: Zou, Bo, et al.
Published: (2024)
SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces
by: Wu, Guande, et al.
Published: (2025)
by: Wu, Guande, et al.
Published: (2025)
Distilling Vision Transformers for Distortion-Robust Representation Learning
by: Alexis, Konstantinos, et al.
Published: (2026)
by: Alexis, Konstantinos, et al.
Published: (2026)
Knowledge Distillation in Vision Transformers: A Critical Review
by: Habib, Gousia, et al.
Published: (2023)
by: Habib, Gousia, et al.
Published: (2023)
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
by: Wu, Size, et al.
Published: (2023)
by: Wu, Size, et al.
Published: (2023)
Revisit Event Generation Model: Self-Supervised Learning of Event-to-Video Reconstruction with Implicit Neural Representations
by: Wang, Zipeng, et al.
Published: (2024)
by: Wang, Zipeng, et al.
Published: (2024)
Depth-Wise Convolutions in Vision Transformers for Efficient Training on Small Datasets
by: Zhang, Tianxiao, et al.
Published: (2024)
by: Zhang, Tianxiao, et al.
Published: (2024)
Comprehensive Survey of Model Compression and Speed up for Vision Transformers
by: Chen, Feiyang, et al.
Published: (2024)
by: Chen, Feiyang, et al.
Published: (2024)
Learning Spatial Decay for Vision Transformers
by: Mao, Yuxin, et al.
Published: (2025)
by: Mao, Yuxin, et al.
Published: (2025)
Scene-Agnostic Traversability Labeling and Estimation via a Multimodal Self-supervised Framework
by: Fang, Zipeng, et al.
Published: (2025)
by: Fang, Zipeng, et al.
Published: (2025)
YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models
by: Chen, Ting, et al.
Published: (2026)
by: Chen, Ting, et al.
Published: (2026)
HyperDiT: Hyper-Connected Transformers for High-Fidelity Pixel-Space Diffusion
by: He, Yu, et al.
Published: (2026)
by: He, Yu, et al.
Published: (2026)
Human-like Object Grouping in Self-supervised Vision Transformers
by: Adeli, Hossein, et al.
Published: (2026)
by: Adeli, Hossein, et al.
Published: (2026)
P2Seg: Pointly-supervised Segmentation via Mutual Distillation
by: Wang, Zipeng, et al.
Published: (2024)
by: Wang, Zipeng, et al.
Published: (2024)
Visual-Advantage On-Policy Distillation for Vision-Language Models
by: Liu, Ruiqi, et al.
Published: (2026)
by: Liu, Ruiqi, et al.
Published: (2026)
Denoising Vision Transformer Autoencoder with Spectral Self-Regularization
by: Xiang, Xunzhi, et al.
Published: (2025)
by: Xiang, Xunzhi, et al.
Published: (2025)
CAS-ViT: Convolutional Additive Self-attention Vision Transformers for Efficient Mobile Applications
by: Zhang, Tianfang, et al.
Published: (2024)
by: Zhang, Tianfang, et al.
Published: (2024)
Object-level Self-Distillation for Vision Pretraining
by: Hızlı, Çağlar, et al.
Published: (2025)
by: Hızlı, Çağlar, et al.
Published: (2025)
EViT: An Eagle Vision Transformer with Bi-Fovea Self-Attention
by: Shi, Yulong, et al.
Published: (2023)
by: Shi, Yulong, et al.
Published: (2023)
On the Faithfulness of Vision Transformer Explanations
by: Wu, Junyi, et al.
Published: (2024)
by: Wu, Junyi, et al.
Published: (2024)
Activation Quantization of Vision Encoders Needs Prefixing Registers
by: Kim, Seunghyeon, et al.
Published: (2025)
by: Kim, Seunghyeon, et al.
Published: (2025)
Self-Supervised Vision Transformer for Enhanced Virtual Clothes Try-On
by: Lu, Lingxiao, et al.
Published: (2024)
by: Lu, Lingxiao, et al.
Published: (2024)
Similar Items
-
EdgeSAM: Prompt-In-the-Loop Distillation for SAM
by: Zhou, Chong, et al.
Published: (2023) -
Vision Transformers Need Registers
by: Darcet, Timothée, et al.
Published: (2023) -
Vision Transformers Need More Than Registers
by: Shi, Cheng, et al.
Published: (2026) -
Elastic Attention Cores for Scalable Vision Transformers
by: Song, Alan Z., et al.
Published: (2026) -
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
by: Zhang, Tianxiao, et al.
Published: (2024)