Leveraging Registers in Vision Transformers for Robust Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Yellapragada, Srikar, Thopalli, Kowshik, Narayanaswamy, Vivek, Sakla, Wesam, Liu, Yang, Mubarka, Yamen, Samaras, Dimitris, Thiagarajan, Jayaraman J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Use of Anchoring for Training Vision Models
by: Narayanaswamy, Vivek, et al.
Published: (2024)
by: Narayanaswamy, Vivek, et al.
Published: (2024)
DECIDER: Leveraging Foundation Model Priors for Improved Model Failure Detection and Explanation
by: Subramanyam, Rakshith, et al.
Published: (2024)
by: Subramanyam, Rakshith, et al.
Published: (2024)
Speeding Up Image Classifiers with Little Companions
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
by: Kulkarni, Akshay, et al.
Published: (2025)
by: Kulkarni, Akshay, et al.
Published: (2025)
Improving Robustness In Sparse Autoencoders via Masked Regularization
by: Narayanaswamy, Vivek, et al.
Published: (2026)
by: Narayanaswamy, Vivek, et al.
Published: (2026)
PathSegDiff: Pathology Segmentation using Diffusion model representations
by: Danisetty, Sachin Kumar, et al.
Published: (2025)
by: Danisetty, Sachin Kumar, et al.
Published: (2025)
$\infty$-Brush: Controllable Large Image Synthesis with Diffusion Models in Infinite Dimensions
by: Le, Minh-Quan, et al.
Published: (2024)
by: Le, Minh-Quan, et al.
Published: (2024)
GECKO: Gigapixel Vision-Concept Contrastive Pretraining in Histopathology
by: Kapse, Saarthak, et al.
Published: (2025)
by: Kapse, Saarthak, et al.
Published: (2025)
Learned representation-guided diffusion models for large-image generation
by: Graikos, Alexandros, et al.
Published: (2023)
by: Graikos, Alexandros, et al.
Published: (2023)
ZoomLDM: Latent Diffusion Model for multi-scale image generation
by: Yellapragada, Srikar, et al.
Published: (2024)
by: Yellapragada, Srikar, et al.
Published: (2024)
Gen-SIS: Generative Self-augmentation Improves Self-supervised Learning
by: Belagali, Varun, et al.
Published: (2024)
by: Belagali, Varun, et al.
Published: (2024)
CDG-MAE: Learning Correspondences from Diffusion Generated Views
by: Belagali, Varun, et al.
Published: (2025)
by: Belagali, Varun, et al.
Published: (2025)
Pathology Image Compression with Pre-trained Autoencoders
by: Yellapragada, Srikar, et al.
Published: (2025)
by: Yellapragada, Srikar, et al.
Published: (2025)
Low-Rank Head Avatar Personalization with Registers
by: Chakkera, Sai Tanmay Reddy, et al.
Published: (2025)
by: Chakkera, Sai Tanmay Reddy, et al.
Published: (2025)
LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images
by: Flora, James, et al.
Published: (2026)
by: Flora, James, et al.
Published: (2026)
Self-supervised co-salient object detection via feature correspondence at multiple scales
by: Chakraborty, Souradeep, et al.
Published: (2024)
by: Chakraborty, Souradeep, et al.
Published: (2024)
Vision Transformers Need Registers
by: Darcet, Timothée, et al.
Published: (2023)
by: Darcet, Timothée, et al.
Published: (2023)
Vision Transformers Need More Than Registers
by: Shi, Cheng, et al.
Published: (2026)
by: Shi, Cheng, et al.
Published: (2026)
ProtAlign: Contrastive learning paradigm for Sequence and structure alignment
by: Ranganath, Aditya, et al.
Published: (2026)
by: Ranganath, Aditya, et al.
Published: (2026)
Vision Transformers with Self-Distilled Registers
by: Chen, Yinjie, et al.
Published: (2025)
by: Chen, Yinjie, et al.
Published: (2025)
TopoDiffusionNet: A Topology-aware Diffusion Model
by: Gupta, Saumya, et al.
Published: (2024)
by: Gupta, Saumya, et al.
Published: (2024)
Weighting Pseudo-Labels via High-Activation Feature Index Similarity and Object Detection for Semi-Supervised Segmentation
by: Howlader, Prantik, et al.
Published: (2024)
by: Howlader, Prantik, et al.
Published: (2024)
Assessing Sample Quality via the Latent Space of Generative Models
by: Xu, Jingyi, et al.
Published: (2024)
by: Xu, Jingyi, et al.
Published: (2024)
One Attention, One Scale: Phase-Aligned Rotary Positional Embeddings for Mixed-Resolution Diffusion Transformer
by: Wu, Haoyu, et al.
Published: (2025)
by: Wu, Haoyu, et al.
Published: (2025)
Enhancing Accuracy and Parameter-Efficiency of Neural Representations for Network Parameterization
by: Choi, Hongjun, et al.
Published: (2024)
by: Choi, Hongjun, et al.
Published: (2024)
TICON: A Slide-Level Tile Contextualizer for Histopathology Representation Learning
by: Belagali, Varun, et al.
Published: (2025)
by: Belagali, Varun, et al.
Published: (2025)
MI-NeRF: Learning a Single Face NeRF from Multiple Identities
by: Chatziagapi, Aggelina, et al.
Published: (2024)
by: Chatziagapi, Aggelina, et al.
Published: (2024)
MIGS: Multi-Identity Gaussian Splatting via Tensor Decomposition
by: Chatziagapi, Aggelina, et al.
Published: (2024)
by: Chatziagapi, Aggelina, et al.
Published: (2024)
The Anatomy of Uncertainty in LLMs
by: Taparia, Aditya, et al.
Published: (2026)
by: Taparia, Aditya, et al.
Published: (2026)
`Eyes of a Hawk and Ears of a Fox': Part Prototype Network for Generalized Zero-Shot Learning
by: Feinglass, Joshua, et al.
Published: (2024)
by: Feinglass, Joshua, et al.
Published: (2024)
Multi-view Gaze Target Estimation
by: Miao, Qiaomu, et al.
Published: (2025)
by: Miao, Qiaomu, et al.
Published: (2025)
Fast constrained sampling in pre-trained diffusion models
by: Graikos, Alexandros, et al.
Published: (2024)
by: Graikos, Alexandros, et al.
Published: (2024)
Learning 3D Reconstruction with Priors in Test Time
by: Zhou, Lei, et al.
Published: (2026)
by: Zhou, Lei, et al.
Published: (2026)
Beyond Pixels: Semi-Supervised Semantic Segmentation with a Multi-scale Patch-based Multi-Label Classifier
by: Howlader, Prantik, et al.
Published: (2024)
by: Howlader, Prantik, et al.
Published: (2024)
Importance-Based Token Merging for Efficient Image and Video Generation
by: Wu, Haoyu, et al.
Published: (2024)
by: Wu, Haoyu, et al.
Published: (2024)
JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation
by: Chakkera, Sai Tanmay Reddy, et al.
Published: (2024)
by: Chakkera, Sai Tanmay Reddy, et al.
Published: (2024)
AV-Flow: Transforming Text to Audio-Visual Human-like Interactions
by: Chatziagapi, Aggelina, et al.
Published: (2025)
by: Chatziagapi, Aggelina, et al.
Published: (2025)
Vision Transformers Don't Need Trained Registers
by: Jiang, Nick, et al.
Published: (2025)
by: Jiang, Nick, et al.
Published: (2025)
Rig3DGS: Creating Controllable Portraits from Casual Monocular Videos
by: Rivero, Alfredo, et al.
Published: (2024)
by: Rivero, Alfredo, et al.
Published: (2024)
What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards
by: Le, Minh-Quan, et al.
Published: (2025)
by: Le, Minh-Quan, et al.
Published: (2025)
Similar Items
-
On the Use of Anchoring for Training Vision Models
by: Narayanaswamy, Vivek, et al.
Published: (2024) -
DECIDER: Leveraging Foundation Model Priors for Improved Model Failure Detection and Explanation
by: Subramanyam, Rakshith, et al.
Published: (2024) -
Speeding Up Image Classifiers with Little Companions
by: Liu, Yang, et al.
Published: (2024) -
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
by: Kulkarni, Akshay, et al.
Published: (2025) -
Improving Robustness In Sparse Autoencoders via Masked Regularization
by: Narayanaswamy, Vivek, et al.
Published: (2026)