Vision Transformers with Self-Distilled Registers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yinjie, Yan, Zipeng, Zhou, Chong, Dai, Bo, Luo, Andrew F. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Vision Transformers Need Registers
von: Darcet, Timothée, et al.
Veröffentlicht: (2023)
von: Darcet, Timothée, et al.
Veröffentlicht: (2023)
EdgeSAM: Prompt-In-the-Loop Distillation for SAM
von: Zhou, Chong, et al.
Veröffentlicht: (2023)
von: Zhou, Chong, et al.
Veröffentlicht: (2023)
Vision Transformers Need More Than Registers
von: Shi, Cheng, et al.
Veröffentlicht: (2026)
von: Shi, Cheng, et al.
Veröffentlicht: (2026)
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
von: Zhang, Tianxiao, et al.
Veröffentlicht: (2024)
von: Zhang, Tianxiao, et al.
Veröffentlicht: (2024)
Elastic Attention Cores for Scalable Vision Transformers
von: Song, Alan Z., et al.
Veröffentlicht: (2026)
von: Song, Alan Z., et al.
Veröffentlicht: (2026)
Vision Transformers Don't Need Trained Registers
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
von: Jiang, Nick, et al.
Veröffentlicht: (2025)
Leveraging Registers in Vision Transformers for Robust Adaptation
von: Yellapragada, Srikar, et al.
Veröffentlicht: (2025)
von: Yellapragada, Srikar, et al.
Veröffentlicht: (2025)
Video Set Distillation: Information Diversification and Temporal Densification
von: Zhao, Yinjie, et al.
Veröffentlicht: (2024)
von: Zhao, Yinjie, et al.
Veröffentlicht: (2024)
Mamba-R: Vision Mamba ALSO Needs Registers
von: Wang, Feng, et al.
Veröffentlicht: (2024)
von: Wang, Feng, et al.
Veröffentlicht: (2024)
Distill-SODA: Distilling Self-Supervised Vision Transformer for Source-Free Open-Set Domain Adaptation in Computational Pathology
von: Vray, Guillaume, et al.
Veröffentlicht: (2023)
von: Vray, Guillaume, et al.
Veröffentlicht: (2023)
FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
von: Wang, Zipeng, et al.
Veröffentlicht: (2025)
von: Wang, Zipeng, et al.
Veröffentlicht: (2025)
Distillation Dynamics: Towards Understanding Feature-Based Distillation in Vision Transformers
von: Tian, Huiyuan, et al.
Veröffentlicht: (2025)
von: Tian, Huiyuan, et al.
Veröffentlicht: (2025)
Generalizing Vision-Language Models with Dedicated Prompt Guidance
von: Li, Xinyao, et al.
Veröffentlicht: (2025)
von: Li, Xinyao, et al.
Veröffentlicht: (2025)
Zero-Ablation Overstates Register Content Dependence in DINO Vision Transformers
von: Parodi, Felipe, et al.
Veröffentlicht: (2026)
von: Parodi, Felipe, et al.
Veröffentlicht: (2026)
Registers Matter for Pixel-Space Diffusion Transformers
von: Starodubcev, Nikita, et al.
Veröffentlicht: (2026)
von: Starodubcev, Nikita, et al.
Veröffentlicht: (2026)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)
Distilling Vision Transformers for Distortion-Robust Representation Learning
von: Alexis, Konstantinos, et al.
Veröffentlicht: (2026)
von: Alexis, Konstantinos, et al.
Veröffentlicht: (2026)
Knowledge Distillation in Vision Transformers: A Critical Review
von: Habib, Gousia, et al.
Veröffentlicht: (2023)
von: Habib, Gousia, et al.
Veröffentlicht: (2023)
SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces
von: Wu, Guande, et al.
Veröffentlicht: (2025)
von: Wu, Guande, et al.
Veröffentlicht: (2025)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
von: Zou, Bo, et al.
Veröffentlicht: (2024)
von: Zou, Bo, et al.
Veröffentlicht: (2024)
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
von: Wu, Size, et al.
Veröffentlicht: (2023)
von: Wu, Size, et al.
Veröffentlicht: (2023)
Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression
von: Ye, Hancheng, et al.
Veröffentlicht: (2024)
von: Ye, Hancheng, et al.
Veröffentlicht: (2024)
Dynamic-Aware Video Distillation: Optimizing Temporal Resolution Based on Video Semantics
von: Zhao, Yinjie, et al.
Veröffentlicht: (2025)
von: Zhao, Yinjie, et al.
Veröffentlicht: (2025)
Depth-Wise Convolutions in Vision Transformers for Efficient Training on Small Datasets
von: Zhang, Tianxiao, et al.
Veröffentlicht: (2024)
von: Zhang, Tianxiao, et al.
Veröffentlicht: (2024)
Comprehensive Survey of Model Compression and Speed up for Vision Transformers
von: Chen, Feiyang, et al.
Veröffentlicht: (2024)
von: Chen, Feiyang, et al.
Veröffentlicht: (2024)
YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models
von: Chen, Ting, et al.
Veröffentlicht: (2026)
von: Chen, Ting, et al.
Veröffentlicht: (2026)
Learning Spatial Decay for Vision Transformers
von: Mao, Yuxin, et al.
Veröffentlicht: (2025)
von: Mao, Yuxin, et al.
Veröffentlicht: (2025)
Revisit Event Generation Model: Self-Supervised Learning of Event-to-Video Reconstruction with Implicit Neural Representations
von: Wang, Zipeng, et al.
Veröffentlicht: (2024)
von: Wang, Zipeng, et al.
Veröffentlicht: (2024)
Denoising Vision Transformer Autoencoder with Spectral Self-Regularization
von: Xiang, Xunzhi, et al.
Veröffentlicht: (2025)
von: Xiang, Xunzhi, et al.
Veröffentlicht: (2025)
Visual-Advantage On-Policy Distillation for Vision-Language Models
von: Liu, Ruiqi, et al.
Veröffentlicht: (2026)
von: Liu, Ruiqi, et al.
Veröffentlicht: (2026)
EViT: An Eagle Vision Transformer with Bi-Fovea Self-Attention
von: Shi, Yulong, et al.
Veröffentlicht: (2023)
von: Shi, Yulong, et al.
Veröffentlicht: (2023)
CAS-ViT: Convolutional Additive Self-attention Vision Transformers for Efficient Mobile Applications
von: Zhang, Tianfang, et al.
Veröffentlicht: (2024)
von: Zhang, Tianfang, et al.
Veröffentlicht: (2024)
Scene-Agnostic Traversability Labeling and Estimation via a Multimodal Self-supervised Framework
von: Fang, Zipeng, et al.
Veröffentlicht: (2025)
von: Fang, Zipeng, et al.
Veröffentlicht: (2025)
On the Faithfulness of Vision Transformer Explanations
von: Wu, Junyi, et al.
Veröffentlicht: (2024)
von: Wu, Junyi, et al.
Veröffentlicht: (2024)
HyperDiT: Hyper-Connected Transformers for High-Fidelity Pixel-Space Diffusion
von: He, Yu, et al.
Veröffentlicht: (2026)
von: He, Yu, et al.
Veröffentlicht: (2026)
Self-Supervised Vision Transformer for Enhanced Virtual Clothes Try-On
von: Lu, Lingxiao, et al.
Veröffentlicht: (2024)
von: Lu, Lingxiao, et al.
Veröffentlicht: (2024)
Object-level Self-Distillation for Vision Pretraining
von: Hızlı, Çağlar, et al.
Veröffentlicht: (2025)
von: Hızlı, Çağlar, et al.
Veröffentlicht: (2025)
P2Seg: Pointly-supervised Segmentation via Mutual Distillation
von: Wang, Zipeng, et al.
Veröffentlicht: (2024)
von: Wang, Zipeng, et al.
Veröffentlicht: (2024)
Human-like Object Grouping in Self-supervised Vision Transformers
von: Adeli, Hossein, et al.
Veröffentlicht: (2026)
von: Adeli, Hossein, et al.
Veröffentlicht: (2026)
Analyzing Local Representations of Self-supervised Vision Transformers
von: Vanyan, Ani, et al.
Veröffentlicht: (2023)
von: Vanyan, Ani, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Vision Transformers Need Registers
von: Darcet, Timothée, et al.
Veröffentlicht: (2023) -
EdgeSAM: Prompt-In-the-Loop Distillation for SAM
von: Zhou, Chong, et al.
Veröffentlicht: (2023) -
Vision Transformers Need More Than Registers
von: Shi, Cheng, et al.
Veröffentlicht: (2026) -
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
von: Zhang, Tianxiao, et al.
Veröffentlicht: (2024) -
Elastic Attention Cores for Scalable Vision Transformers
von: Song, Alan Z., et al.
Veröffentlicht: (2026)