UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qiu, Congpei, Hu, Zhaoyu, Ke, Wei, Tian, Zhuotao, Wu, Yanhao, Zhang, Tong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Refining Datapath for Microscaling ViTs
von: Xiao, Can, et al.
Veröffentlicht: (2025)
von: Xiao, Can, et al.
Veröffentlicht: (2025)
Refining CLIP's Spatial Awareness: A Visual-Centric Perspective
von: Qiu, Congpei, et al.
Veröffentlicht: (2025)
von: Qiu, Congpei, et al.
Veröffentlicht: (2025)
ViT Registers and Fractal ViT
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2026)
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2026)
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
von: Shah, Arya, et al.
Veröffentlicht: (2025)
von: Shah, Arya, et al.
Veröffentlicht: (2025)
CLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-Training Quantization of ViTs
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
von: Zhong, Yunshan, et al.
Veröffentlicht: (2023)
von: Zhong, Yunshan, et al.
Veröffentlicht: (2023)
Register and [CLS] tokens yield a decoupling of local and global features in large ViTs
von: Lappe, Alexander, et al.
Veröffentlicht: (2025)
von: Lappe, Alexander, et al.
Veröffentlicht: (2025)
U-REPA: Aligning Diffusion U-Nets to ViTs
von: Tian, Yuchuan, et al.
Veröffentlicht: (2025)
von: Tian, Yuchuan, et al.
Veröffentlicht: (2025)
Colinearity Decay: Training Quantization-Friendly ViTs with Outlier Decay
von: Tong, Jin, et al.
Veröffentlicht: (2026)
von: Tong, Jin, et al.
Veröffentlicht: (2026)
Mitigating Object Dependencies: Improving Point Cloud Self-Supervised Learning through Object Exchange
von: Wu, Yanhao, et al.
Veröffentlicht: (2024)
von: Wu, Yanhao, et al.
Veröffentlicht: (2024)
Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
von: Chen, Lu, et al.
Veröffentlicht: (2025)
von: Chen, Lu, et al.
Veröffentlicht: (2025)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
von: Chattopadhyay, Nandish, et al.
Veröffentlicht: (2026)
von: Chattopadhyay, Nandish, et al.
Veröffentlicht: (2026)
ViTs: Teaching Machines to See Time Series Anomalies Like Human Experts
von: Wang, Zexin, et al.
Veröffentlicht: (2025)
von: Wang, Zexin, et al.
Veröffentlicht: (2025)
Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
von: Siméoni, Oriane, et al.
Veröffentlicht: (2023)
von: Siméoni, Oriane, et al.
Veröffentlicht: (2023)
VIVID-Med: LLM-Supervised Structured Pretraining for Deployable Medical ViTs
von: Wang, Xiyao, et al.
Veröffentlicht: (2026)
von: Wang, Xiyao, et al.
Veröffentlicht: (2026)
Elastic ViTs from Pretrained Models without Retraining
von: Simoncini, Walter, et al.
Veröffentlicht: (2025)
von: Simoncini, Walter, et al.
Veröffentlicht: (2025)
Harnessing the Computation Redundancy in ViTs to Boost Adversarial Transferability
von: Liu, Jiani, et al.
Veröffentlicht: (2025)
von: Liu, Jiani, et al.
Veröffentlicht: (2025)
Training-Free Acceleration of ViTs with Delayed Spatial Merging
von: Heo, Jung Hwan, et al.
Veröffentlicht: (2023)
von: Heo, Jung Hwan, et al.
Veröffentlicht: (2023)
Fusing Pretrained ViTs with TCNet for Enhanced EEG Regression
von: Modesitt, Eric, et al.
Veröffentlicht: (2024)
von: Modesitt, Eric, et al.
Veröffentlicht: (2024)
Pretrained ViTs Yield Versatile Representations For Medical Images
von: Matsoukas, Christos, et al.
Veröffentlicht: (2023)
von: Matsoukas, Christos, et al.
Veröffentlicht: (2023)
Octic Vision Transformers: Quicker ViTs Through Equivariance
von: Nordström, David, et al.
Veröffentlicht: (2025)
von: Nordström, David, et al.
Veröffentlicht: (2025)
Parameter Efficient Fine-tuning of Self-supervised ViTs without Catastrophic Forgetting
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2024)
von: Bafghi, Reza Akbarian, et al.
Veröffentlicht: (2024)
Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP
von: Balasubramanian, Sriram, et al.
Veröffentlicht: (2024)
von: Balasubramanian, Sriram, et al.
Veröffentlicht: (2024)
Let ViT Speak: Generative Language-Image Pre-training
von: Fang, Yan, et al.
Veröffentlicht: (2026)
von: Fang, Yan, et al.
Veröffentlicht: (2026)
ConcatPlexer: Additional Dim1 Batching for Faster ViTs
von: Han, Donghoon, et al.
Veröffentlicht: (2023)
von: Han, Donghoon, et al.
Veröffentlicht: (2023)
Token Cropr: Faster ViTs for Quite a Few Tasks
von: Bergner, Benjamin, et al.
Veröffentlicht: (2024)
von: Bergner, Benjamin, et al.
Veröffentlicht: (2024)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
von: Hwang, Dongyoon, et al.
Veröffentlicht: (2024)
von: Hwang, Dongyoon, et al.
Veröffentlicht: (2024)
Which Direction to Choose? An Analysis on the Representation Power of Self-Supervised ViTs in Downstream Tasks
von: Kaltampanidis, Yannis, et al.
Veröffentlicht: (2025)
von: Kaltampanidis, Yannis, et al.
Veröffentlicht: (2025)
TAP-ViTs: Task-Adaptive Pruning for On-Device Deployment of Vision Transformers
von: Wang, Zhibo, et al.
Veröffentlicht: (2026)
von: Wang, Zhibo, et al.
Veröffentlicht: (2026)
Communication Efficient Split Learning of ViTs with Attention-based Double Compression
von: Alvetreti, Federico, et al.
Veröffentlicht: (2025)
von: Alvetreti, Federico, et al.
Veröffentlicht: (2025)
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
von: Kim, Donghyun, et al.
Veröffentlicht: (2024)
von: Kim, Donghyun, et al.
Veröffentlicht: (2024)
EdgeCrafter: Compact ViTs for Edge Dense Prediction via Task-Specialized Distillation
von: Liu, Longfei, et al.
Veröffentlicht: (2026)
von: Liu, Longfei, et al.
Veröffentlicht: (2026)
ViTs are Everywhere: A Comprehensive Study Showcasing Vision Transformers in Different Domain
von: Mia, Md Sohag, et al.
Veröffentlicht: (2023)
von: Mia, Md Sohag, et al.
Veröffentlicht: (2023)
Exploring the Synergies of Hybrid CNNs and ViTs Architectures for Computer Vision: A survey
von: Yunusa, Haruna, et al.
Veröffentlicht: (2024)
von: Yunusa, Haruna, et al.
Veröffentlicht: (2024)
Causality $\neq$ Decodability, and Vice Versa: Lessons from Interpreting Counting ViTs
von: Huang, Lianghuan, et al.
Veröffentlicht: (2025)
von: Huang, Lianghuan, et al.
Veröffentlicht: (2025)
Provable Sparse Inversion and Token Relabel Enhanced One-shot Federated Learning with ViTs
von: Shen, Li, et al.
Veröffentlicht: (2026)
von: Shen, Li, et al.
Veröffentlicht: (2026)
AFIDAF: Alternating Fourier and Image Domain Adaptive Filters as an Efficient Alternative to Attention in ViTs
von: Zheng, Yunling, et al.
Veröffentlicht: (2024)
von: Zheng, Yunling, et al.
Veröffentlicht: (2024)
Generating Multimodal Driving Scenes via Next-Scene Prediction
von: Wu, Yanhao, et al.
Veröffentlicht: (2025)
von: Wu, Yanhao, et al.
Veröffentlicht: (2025)
How to train your ViT for OOD Detection
von: Mueller, Maximilian, et al.
Veröffentlicht: (2024)
von: Mueller, Maximilian, et al.
Veröffentlicht: (2024)
Modulating CNN Features with Pre-Trained ViT Representations for Open-Vocabulary Object Detection
von: Gao, Xiangyu, et al.
Veröffentlicht: (2025)
von: Gao, Xiangyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Refining Datapath for Microscaling ViTs
von: Xiao, Can, et al.
Veröffentlicht: (2025) -
Refining CLIP's Spatial Awareness: A Visual-Centric Perspective
von: Qiu, Congpei, et al.
Veröffentlicht: (2025) -
ViT Registers and Fractal ViT
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2026) -
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
von: Shah, Arya, et al.
Veröffentlicht: (2025) -
CLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-Training Quantization of ViTs
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)