Layer by layer, module by module: Choose both for optimal OOD probing of ViT
Fuente:
arXiv
Saved in:
| Main Authors: | Odonnat, Ambroise, Feofanov, Vasilii, Chapel, Laetitia, Tavenard, Romain, Redko, Ievgen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision Transformer Finetuning Benefits from Non-Smooth Components
by: Odonnat, Ambroise, et al.
Published: (2026)
by: Odonnat, Ambroise, et al.
Published: (2026)
Leveraging Ensemble Diversity for Robust Self-Training in the Presence of Sample Selection Bias
by: Odonnat, Ambroise, et al.
Published: (2023)
by: Odonnat, Ambroise, et al.
Published: (2023)
Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers
by: Roschmann, Simon, et al.
Published: (2025)
by: Roschmann, Simon, et al.
Published: (2025)
Leveraging Gradients for Unsupervised Accuracy Estimation under Distribution Shift
by: Xie, Renchunzi, et al.
Published: (2024)
by: Xie, Renchunzi, et al.
Published: (2024)
Analysing Multi-Task Regression via Random Matrix Theory with Application to Time Series Forecasting
by: Ilbert, Romain, et al.
Published: (2024)
by: Ilbert, Romain, et al.
Published: (2024)
SAMformer: Unlocking the Potential of Transformers in Time Series Forecasting with Sharpness-Aware Minimization and Channel-Wise Attention
by: Ilbert, Romain, et al.
Published: (2024)
by: Ilbert, Romain, et al.
Published: (2024)
How to train your ViT for OOD Detection
by: Mueller, Maximilian, et al.
Published: (2024)
by: Mueller, Maximilian, et al.
Published: (2024)
Differentiable Generalized Sliced Wasserstein Plans
by: Chapel, Laetitia, et al.
Published: (2025)
by: Chapel, Laetitia, et al.
Published: (2025)
User-friendly Foundation Model Adapters for Multivariate Time Series Classification
by: Feofanov, Vasilii, et al.
Published: (2024)
by: Feofanov, Vasilii, et al.
Published: (2024)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
by: Chattopadhyay, Nandish, et al.
Published: (2026)
by: Chattopadhyay, Nandish, et al.
Published: (2026)
ViTCAE: ViT-based Class-conditioned Autoencoder
by: Jebraeeli, Vahid, et al.
Published: (2025)
by: Jebraeeli, Vahid, et al.
Published: (2025)
CauKer: Classification Time Series Foundation Models Can Be Pretrained on Synthetic Data
by: Xie, Shifeng, et al.
Published: (2025)
by: Xie, Shifeng, et al.
Published: (2025)
HydraViT: Stacking Heads for a Scalable ViT
by: Haberer, Janek, et al.
Published: (2024)
by: Haberer, Janek, et al.
Published: (2024)
Leveraging Generic Time Series Foundation Models for EEG Classification
by: Gnassounou, Théo, et al.
Published: (2025)
by: Gnassounou, Théo, et al.
Published: (2025)
MantisV2: Closing the Zero-Shot Gap in Time Series Classification with Synthetic Data and Test-Time Strategies
by: Feofanov, Vasilii, et al.
Published: (2026)
by: Feofanov, Vasilii, et al.
Published: (2026)
Intriguing Frequency Interpretation of Adversarial Robustness for CNNs and ViTs
by: Chen, Lu, et al.
Published: (2025)
by: Chen, Lu, et al.
Published: (2025)
$Δ\mathrm{Energy}$: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization
by: Zhu, Lin, et al.
Published: (2025)
by: Zhu, Lin, et al.
Published: (2025)
MANO: Exploiting Matrix Norm for Unsupervised Accuracy Estimation Under Distribution Shifts
by: Xie, Renchunzi, et al.
Published: (2024)
by: Xie, Renchunzi, et al.
Published: (2024)
ProtoS-ViT: Visual foundation models for sparse self-explainable classifications
by: Turbé, Hugues, et al.
Published: (2024)
by: Turbé, Hugues, et al.
Published: (2024)
Exploiting Lightweight Hierarchical ViT and Dynamic Framework for Efficient Visual Tracking
by: Kang, Ben, et al.
Published: (2025)
by: Kang, Ben, et al.
Published: (2025)
Token Cropr: Faster ViTs for Quite a Few Tasks
by: Bergner, Benjamin, et al.
Published: (2024)
by: Bergner, Benjamin, et al.
Published: (2024)
CubistMerge: Spatial-Preserving Token Merging For Diverse ViT Backbones
by: Gong, Wenyi, et al.
Published: (2025)
by: Gong, Wenyi, et al.
Published: (2025)
LL-ViT: Edge Deployable Vision Transformers with Look Up Table Neurons
by: Nag, Shashank, et al.
Published: (2025)
by: Nag, Shashank, et al.
Published: (2025)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
by: Hwang, Dongyoon, et al.
Published: (2024)
by: Hwang, Dongyoon, et al.
Published: (2024)
Exploring the Synergies of Hybrid CNNs and ViTs Architectures for Computer Vision: A survey
by: Yunusa, Haruna, et al.
Published: (2024)
by: Yunusa, Haruna, et al.
Published: (2024)
Causality $\neq$ Decodability, and Vice Versa: Lessons from Interpreting Counting ViTs
by: Huang, Lianghuan, et al.
Published: (2025)
by: Huang, Lianghuan, et al.
Published: (2025)
Training-Free Acceleration of ViTs with Delayed Spatial Merging
by: Heo, Jung Hwan, et al.
Published: (2023)
by: Heo, Jung Hwan, et al.
Published: (2023)
Octic Vision Transformers: Quicker ViTs Through Equivariance
by: Nordström, David, et al.
Published: (2025)
by: Nordström, David, et al.
Published: (2025)
Large Language Models as Markov Chains
by: Zekri, Oussama, et al.
Published: (2024)
by: Zekri, Oussama, et al.
Published: (2024)
Register and [CLS] tokens yield a decoupling of local and global features in large ViTs
by: Lappe, Alexander, et al.
Published: (2025)
by: Lappe, Alexander, et al.
Published: (2025)
ChAda-ViT : Channel Adaptive Attention for Joint Representation Learning of Heterogeneous Microscopy Images
by: Bourriez, Nicolas, et al.
Published: (2023)
by: Bourriez, Nicolas, et al.
Published: (2023)
GrowTAS: Progressive Expansion from Small to Large Subnets for Efficient ViT Architecture Search
by: Lee, Hyunju, et al.
Published: (2025)
by: Lee, Hyunju, et al.
Published: (2025)
ViT-MUL: A Baseline Study on Recent Machine Unlearning Methods Applied to Vision Transformers
by: Cho, Ikhyun, et al.
Published: (2024)
by: Cho, Ikhyun, et al.
Published: (2024)
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
by: Tonmoy, Moshiur Rahman, et al.
Published: (2025)
by: Tonmoy, Moshiur Rahman, et al.
Published: (2025)
Triamese-ViT: A 3D-Aware Method for Robust Brain Age Estimation from MRIs
by: Zhang, Zhaonian, et al.
Published: (2024)
by: Zhang, Zhaonian, et al.
Published: (2024)
ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers
by: Sinha, Sanchit, et al.
Published: (2025)
by: Sinha, Sanchit, et al.
Published: (2025)
Tex-ViT: A Generalizable, Robust, Texture-based dual-branch cross-attention deepfake detector
by: Dagar, Deepak, et al.
Published: (2024)
by: Dagar, Deepak, et al.
Published: (2024)
ConcatPlexer: Additional Dim1 Batching for Faster ViTs
by: Han, Donghoon, et al.
Published: (2023)
by: Han, Donghoon, et al.
Published: (2023)
ViT Enhanced Privacy-Preserving Secure Medical Data Sharing and Classification
by: Amin, Al, et al.
Published: (2024)
by: Amin, Al, et al.
Published: (2024)
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
by: Li, Zhengang, et al.
Published: (2024)
by: Li, Zhengang, et al.
Published: (2024)
Similar Items
-
Vision Transformer Finetuning Benefits from Non-Smooth Components
by: Odonnat, Ambroise, et al.
Published: (2026) -
Leveraging Ensemble Diversity for Robust Self-Training in the Presence of Sample Selection Bias
by: Odonnat, Ambroise, et al.
Published: (2023) -
Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers
by: Roschmann, Simon, et al.
Published: (2025) -
Leveraging Gradients for Unsupervised Accuracy Estimation under Distribution Shift
by: Xie, Renchunzi, et al.
Published: (2024) -
Analysing Multi-Task Regression via Random Matrix Theory with Application to Time Series Forecasting
by: Ilbert, Romain, et al.
Published: (2024)