Register and [CLS] tokens yield a decoupling of local and global features in large ViTs
Fuente:
arXiv
Guardado en:
| Autores principales: | Lappe, Alexander, Giese, Martin A. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Another BRIXEL in the Wall: Towards Cheaper Dense Features
por: Lappe, Alexander, et al.
Publicado: (2025)
por: Lappe, Alexander, et al.
Publicado: (2025)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
por: Chattopadhyay, Nandish, et al.
Publicado: (2026)
por: Chattopadhyay, Nandish, et al.
Publicado: (2026)
Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
por: Chen, Lu, et al.
Publicado: (2025)
por: Chen, Lu, et al.
Publicado: (2025)
Token Cropr: Faster ViTs for Quite a Few Tasks
por: Bergner, Benjamin, et al.
Publicado: (2024)
por: Bergner, Benjamin, et al.
Publicado: (2024)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
por: Hwang, Dongyoon, et al.
Publicado: (2024)
por: Hwang, Dongyoon, et al.
Publicado: (2024)
Octic Vision Transformers: Quicker ViTs Through Equivariance
por: Nordström, David, et al.
Publicado: (2025)
por: Nordström, David, et al.
Publicado: (2025)
Training-Free Acceleration of ViTs with Delayed Spatial Merging
por: Heo, Jung Hwan, et al.
Publicado: (2023)
por: Heo, Jung Hwan, et al.
Publicado: (2023)
Causality $\neq$ Decodability, and Vice Versa: Lessons from Interpreting Counting ViTs
por: Huang, Lianghuan, et al.
Publicado: (2025)
por: Huang, Lianghuan, et al.
Publicado: (2025)
Exploring the Synergies of Hybrid CNNs and ViTs Architectures for Computer Vision: A survey
por: Yunusa, Haruna, et al.
Publicado: (2024)
por: Yunusa, Haruna, et al.
Publicado: (2024)
ConcatPlexer: Additional Dim1 Batching for Faster ViTs
por: Han, Donghoon, et al.
Publicado: (2023)
por: Han, Donghoon, et al.
Publicado: (2023)
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs
por: Kim, Donghyun, et al.
Publicado: (2024)
por: Kim, Donghyun, et al.
Publicado: (2024)
Communication Efficient Split Learning of ViTs with Attention-based Double Compression
por: Alvetreti, Federico, et al.
Publicado: (2025)
por: Alvetreti, Federico, et al.
Publicado: (2025)
TAP-ViTs: Task-Adaptive Pruning for On-Device Deployment of Vision Transformers
por: Wang, Zhibo, et al.
Publicado: (2026)
por: Wang, Zhibo, et al.
Publicado: (2026)
Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP
por: Balasubramanian, Sriram, et al.
Publicado: (2024)
por: Balasubramanian, Sriram, et al.
Publicado: (2024)
Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustness
por: Elisha, Yehonatan, et al.
Publicado: (2026)
por: Elisha, Yehonatan, et al.
Publicado: (2026)
ViTCAE: ViT-based Class-conditioned Autoencoder
por: Jebraeeli, Vahid, et al.
Publicado: (2025)
por: Jebraeeli, Vahid, et al.
Publicado: (2025)
UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register
por: Qiu, Congpei, et al.
Publicado: (2026)
por: Qiu, Congpei, et al.
Publicado: (2026)
Parallel Backpropagation for Shared-Feature Visualization
por: Lappe, Alexander, et al.
Publicado: (2024)
por: Lappe, Alexander, et al.
Publicado: (2024)
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
por: Shah, Arya, et al.
Publicado: (2025)
por: Shah, Arya, et al.
Publicado: (2025)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
por: Zhong, Yunshan, et al.
Publicado: (2023)
por: Zhong, Yunshan, et al.
Publicado: (2023)
How to train your ViT for OOD Detection
por: Mueller, Maximilian, et al.
Publicado: (2024)
por: Mueller, Maximilian, et al.
Publicado: (2024)
HydraViT: Stacking Heads for a Scalable ViT
por: Haberer, Janek, et al.
Publicado: (2024)
por: Haberer, Janek, et al.
Publicado: (2024)
U-REPA: Aligning Diffusion U-Nets to ViTs
por: Tian, Yuchuan, et al.
Publicado: (2025)
por: Tian, Yuchuan, et al.
Publicado: (2025)
Elastic ViTs from Pretrained Models without Retraining
por: Simoncini, Walter, et al.
Publicado: (2025)
por: Simoncini, Walter, et al.
Publicado: (2025)
Harnessing the Computation Redundancy in ViTs to Boost Adversarial Transferability
por: Liu, Jiani, et al.
Publicado: (2025)
por: Liu, Jiani, et al.
Publicado: (2025)
Pretrained ViTs Yield Versatile Representations For Medical Images
por: Matsoukas, Christos, et al.
Publicado: (2023)
por: Matsoukas, Christos, et al.
Publicado: (2023)
Colinearity Decay: Training Quantization-Friendly ViTs with Outlier Decay
por: Tong, Jin, et al.
Publicado: (2026)
por: Tong, Jin, et al.
Publicado: (2026)
Exploiting Lightweight Hierarchical ViT and Dynamic Framework for Efficient Visual Tracking
por: Kang, Ben, et al.
Publicado: (2025)
por: Kang, Ben, et al.
Publicado: (2025)
CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
por: Gong, Wenyi, et al.
Publicado: (2025)
por: Gong, Wenyi, et al.
Publicado: (2025)
ProtoS-ViT: Visual foundation models for sparse self-explainable classifications
por: Turbé, Hugues, et al.
Publicado: (2024)
por: Turbé, Hugues, et al.
Publicado: (2024)
LL-ViT: Edge Deployable Vision Transformers with Look Up Table Neurons
por: Nag, Shashank, et al.
Publicado: (2025)
por: Nag, Shashank, et al.
Publicado: (2025)
Layer by layer, module by module: Choose both for optimal OOD probing of ViT
por: Odonnat, Ambroise, et al.
Publicado: (2026)
por: Odonnat, Ambroise, et al.
Publicado: (2026)
Unlocking [CLS] Features for Continual Post-Training
por: Yildirim, Murat Onur, et al.
Publicado: (2025)
por: Yildirim, Murat Onur, et al.
Publicado: (2025)
Unsqueeze [CLS] Bottleneck to Learn Rich Representations
por: Su, Qing, et al.
Publicado: (2024)
por: Su, Qing, et al.
Publicado: (2024)
Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
por: Siméoni, Oriane, et al.
Publicado: (2023)
por: Siméoni, Oriane, et al.
Publicado: (2023)
LookupViT: Compressing visual information to a limited number of tokens
por: Koner, Rajat, et al.
Publicado: (2024)
por: Koner, Rajat, et al.
Publicado: (2024)
CLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-Training Quantization of ViTs
por: Ramachandran, Akshat, et al.
Publicado: (2024)
por: Ramachandran, Akshat, et al.
Publicado: (2024)
GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search
por: Lee, Hyunju, et al.
Publicado: (2025)
por: Lee, Hyunju, et al.
Publicado: (2025)
ChAda-ViT : Channel Adaptive Attention for Joint Representation Learning of Heterogeneous Microscopy Images
por: Bourriez, Nicolas, et al.
Publicado: (2023)
por: Bourriez, Nicolas, et al.
Publicado: (2023)
ViT-MUL: A Baseline Study on Recent Machine Unlearning Methods Applied to Vision Transformers
por: Cho, Ikhyun, et al.
Publicado: (2024)
por: Cho, Ikhyun, et al.
Publicado: (2024)
Ejemplares similares
-
Another BRIXEL in the Wall: Towards Cheaper Dense Features
por: Lappe, Alexander, et al.
Publicado: (2025) -
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
por: Chattopadhyay, Nandish, et al.
Publicado: (2026) -
Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
por: Chen, Lu, et al.
Publicado: (2025) -
Token Cropr: Faster ViTs for Quite a Few Tasks
por: Bergner, Benjamin, et al.
Publicado: (2024) -
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
por: Hwang, Dongyoon, et al.
Publicado: (2024)