Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Morelli, Fabian, Uselis, Arnas, Sonthalia, Ankit, Oh, Seong Joon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the rankability of visual embeddings
von: Sonthalia, Ankit, et al.
Veröffentlicht: (2025)
von: Sonthalia, Ankit, et al.
Veröffentlicht: (2025)
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
von: Koishigarina, Darina, et al.
Veröffentlicht: (2025)
von: Koishigarina, Darina, et al.
Veröffentlicht: (2025)
How can embedding models bind concepts?
von: Uselis, Arnas, et al.
Veröffentlicht: (2026)
von: Uselis, Arnas, et al.
Veröffentlicht: (2026)
Half-Truths Break Similarity-Based Retrieval
von: Kargi, Bora, et al.
Veröffentlicht: (2026)
von: Kargi, Bora, et al.
Veröffentlicht: (2026)
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
von: Uselis, Arnas, et al.
Veröffentlicht: (2026)
von: Uselis, Arnas, et al.
Veröffentlicht: (2026)
Diffusion Classifiers Understand Compositionality, but Conditions Apply
von: Jeong, Yujin, et al.
Veröffentlicht: (2025)
von: Jeong, Yujin, et al.
Veröffentlicht: (2025)
When Do Diffusion Models learn to Generate Multiple Objects?
von: Jeong, Yujin, et al.
Veröffentlicht: (2026)
von: Jeong, Yujin, et al.
Veröffentlicht: (2026)
GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
von: Kim, Nayeong, et al.
Veröffentlicht: (2025)
von: Kim, Nayeong, et al.
Veröffentlicht: (2025)
Interpreting CLIP with Hierarchical Sparse Autoencoders
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2025)
von: Zaigrajew, Vladimir, et al.
Veröffentlicht: (2025)
Semantic-aware Adversarial Fine-tuning for CLIP
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2026)
von: Zhang, Jiacheng, et al.
Veröffentlicht: (2026)
Intermediate Layer Classifiers for OOD generalization
von: Uselis, Arnas, et al.
Veröffentlicht: (2025)
von: Uselis, Arnas, et al.
Veröffentlicht: (2025)
Beyond Semantics: Disentangling Information Scope in Sparse Autoencoders for CLIP
von: Ro, Yusung, et al.
Veröffentlicht: (2026)
von: Ro, Yusung, et al.
Veröffentlicht: (2026)
Matricial Free Energy as a Gaussianizing Regularizer: Enhancing Autoencoders for Gaussian Code Generation
von: Sonthalia, Rishi, et al.
Veröffentlicht: (2025)
von: Sonthalia, Rishi, et al.
Veröffentlicht: (2025)
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)
von: Nguyen, Nghia, et al.
Veröffentlicht: (2024)
Fully Fine-tuned CLIP Models are Efficient Few-Shot Learners
von: Liu, Mushui, et al.
Veröffentlicht: (2024)
von: Liu, Mushui, et al.
Veröffentlicht: (2024)
Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive Learning
von: Qin, Chuan, et al.
Veröffentlicht: (2026)
von: Qin, Chuan, et al.
Veröffentlicht: (2026)
Interpretable and Testable Vision Features via Sparse Autoencoders
von: Stevens, Samuel, et al.
Veröffentlicht: (2025)
von: Stevens, Samuel, et al.
Veröffentlicht: (2025)
Prototypical Contrastive Learning-based CLIP Fine-tuning for Object Re-identification
von: Li, Jiachen, et al.
Veröffentlicht: (2023)
von: Li, Jiachen, et al.
Veröffentlicht: (2023)
Interpretability Transfer from Language to Vision via Sparse Autoencoders
von: Kravets, Alexey, et al.
Veröffentlicht: (2026)
von: Kravets, Alexey, et al.
Veröffentlicht: (2026)
Realistic Unsupervised CLIP Fine-tuning with Universal Entropy Optimization
von: Liang, Jian, et al.
Veröffentlicht: (2023)
von: Liang, Jian, et al.
Veröffentlicht: (2023)
Breaking the Limits of Open-Weight CLIP: An Optimization Framework for Self-supervised Fine-tuning of CLIP
von: Mehta, Anant, et al.
Veröffentlicht: (2026)
von: Mehta, Anant, et al.
Veröffentlicht: (2026)
AD-CLIP: Adapting Domains in Prompt Space Using CLIP
von: Singha, Mainak, et al.
Veröffentlicht: (2023)
von: Singha, Mainak, et al.
Veröffentlicht: (2023)
Causal Interpretation of Sparse Autoencoder Features in Vision
von: Han, Sangyu, et al.
Veröffentlicht: (2025)
von: Han, Sangyu, et al.
Veröffentlicht: (2025)
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
von: Kulkarni, Akshay, et al.
Veröffentlicht: (2025)
von: Kulkarni, Akshay, et al.
Veröffentlicht: (2025)
Mammo-SAE: Interpreting Breast Cancer Concept Learning with Sparse Autoencoders
von: Nakka, Krishna Kanth
Veröffentlicht: (2025)
von: Nakka, Krishna Kanth
Veröffentlicht: (2025)
Sparse-Tuning: Adapting Vision Transformers with Efficient Fine-tuning and Inference
von: Liu, Ting, et al.
Veröffentlicht: (2024)
von: Liu, Ting, et al.
Veröffentlicht: (2024)
Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning
von: Xiao, Junhao, et al.
Veröffentlicht: (2026)
von: Xiao, Junhao, et al.
Veröffentlicht: (2026)
Sparse Autoencoders for Interpretable Medical Image Representation Learning
von: Wesp, Philipp, et al.
Veröffentlicht: (2026)
von: Wesp, Philipp, et al.
Veröffentlicht: (2026)
Fine-tuning a vision-language model for fracture-surface morphology recognition
von: Liu, Quanliang, et al.
Veröffentlicht: (2026)
von: Liu, Quanliang, et al.
Veröffentlicht: (2026)
Fine-tuning CLIP Text Encoders with Two-step Paraphrasing
von: Kim, Hyunjae, et al.
Veröffentlicht: (2024)
von: Kim, Hyunjae, et al.
Veröffentlicht: (2024)
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
von: Bhalla, Usha, et al.
Veröffentlicht: (2024)
von: Bhalla, Usha, et al.
Veröffentlicht: (2024)
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
von: Monsefi, Amin Karimi, et al.
Veröffentlicht: (2024)
Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
von: Thasarathan, Harrish, et al.
Veröffentlicht: (2025)
von: Thasarathan, Harrish, et al.
Veröffentlicht: (2025)
Explorer: Robust Collection of Interactable GUI Elements
von: Chaimalas, Iason, et al.
Veröffentlicht: (2025)
von: Chaimalas, Iason, et al.
Veröffentlicht: (2025)
Robust Fine-tuning of Zero-shot Models via Variance Reduction
von: Zhu, Beier, et al.
Veröffentlicht: (2024)
von: Zhu, Beier, et al.
Veröffentlicht: (2024)
TSCLIP: Robust CLIP Fine-Tuning for Worldwide Cross-Regional Traffic Sign Recognition
von: Zhao, Guoyang, et al.
Veröffentlicht: (2024)
von: Zhao, Guoyang, et al.
Veröffentlicht: (2024)
Perturb and Recover: Fine-tuning for Effective Backdoor Removal from CLIP
von: Singh, Naman Deep, et al.
Veröffentlicht: (2024)
von: Singh, Naman Deep, et al.
Veröffentlicht: (2024)
Does Data Scaling Lead to Visual Compositional Generalization?
von: Uselis, Arnas, et al.
Veröffentlicht: (2025)
von: Uselis, Arnas, et al.
Veröffentlicht: (2025)
Open-Vocabulary X-ray Prohibited Item Detection via Fine-tuning CLIP
von: Lin, Shuyang, et al.
Veröffentlicht: (2024)
von: Lin, Shuyang, et al.
Veröffentlicht: (2024)
ECCV Caption: Correcting False Negatives by Collecting Machine-and-Human-verified Image-Caption Associations for MS-COCO
von: Chun, Sanghyuk, et al.
Veröffentlicht: (2022)
von: Chun, Sanghyuk, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
On the rankability of visual embeddings
von: Sonthalia, Ankit, et al.
Veröffentlicht: (2025) -
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally
von: Koishigarina, Darina, et al.
Veröffentlicht: (2025) -
How can embedding models bind concepts?
von: Uselis, Arnas, et al.
Veröffentlicht: (2026) -
Half-Truths Break Similarity-Based Retrieval
von: Kargi, Bora, et al.
Veröffentlicht: (2026) -
Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models
von: Uselis, Arnas, et al.
Veröffentlicht: (2026)