Interpreting CLIP with Hierarchical Sparse Autoencoders
Fuente:
arXiv
Saved in:
| Main Authors: | Zaigrajew, Vladimir, Baniecki, Hubert, Biecek, Przemyslaw |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SwordBench: Evaluating Orthogonality of Steering Image Representations
by: Zaigrajew, Vladimir, et al.
Published: (2026)
by: Zaigrajew, Vladimir, et al.
Published: (2026)
Adversarial attacks and defenses in explainable artificial intelligence: A survey
by: Baniecki, Hubert, et al.
Published: (2023)
by: Baniecki, Hubert, et al.
Published: (2023)
Birds look like cars: Adversarial analysis of intrinsically interpretable deep learning
by: Baniecki, Hubert, et al.
Published: (2025)
by: Baniecki, Hubert, et al.
Published: (2025)
LINE: LLM-based Iterative Neuron Explanations for Vision Models
by: Zaigrajew, Vladimir, et al.
Published: (2026)
by: Zaigrajew, Vladimir, et al.
Published: (2026)
Red Teaming Models for Hyperspectral Image Analysis Using Explainable AI
by: Zaigrajew, Vladimir, et al.
Published: (2024)
by: Zaigrajew, Vladimir, et al.
Published: (2024)
Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions
by: Baniecki, Hubert, et al.
Published: (2025)
by: Baniecki, Hubert, et al.
Published: (2025)
Red-Teaming Segment Anything Model
by: Jankowski, Krzysztof, et al.
Published: (2024)
by: Jankowski, Krzysztof, et al.
Published: (2024)
Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers
by: Grzywaczewski, Jakub, et al.
Published: (2026)
by: Grzywaczewski, Jakub, et al.
Published: (2026)
Attributions All the Way Down? The Metagame of Interpretability
by: Baniecki, Hubert, et al.
Published: (2026)
by: Baniecki, Hubert, et al.
Published: (2026)
Position: Do Not Explain Vision Models Without Context
by: Tomaszewska, Paulina, et al.
Published: (2024)
by: Tomaszewska, Paulina, et al.
Published: (2024)
Global Counterfactual Directions
by: Sobieski, Bartlomiej, et al.
Published: (2024)
by: Sobieski, Bartlomiej, et al.
Published: (2024)
NormEnsembleXAI: Unveiling the Strengths and Weaknesses of XAI Ensemble Techniques
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
Interpretable machine learning for time-to-event prediction in medicine and healthcare
by: Baniecki, Hubert, et al.
Published: (2023)
by: Baniecki, Hubert, et al.
Published: (2023)
Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
by: Yeung, Calvin, et al.
Published: (2026)
by: Yeung, Calvin, et al.
Published: (2026)
X-ray transferable polyrepresentation learning
by: Hryniewska-Guzik, Weronika, et al.
Published: (2025)
by: Hryniewska-Guzik, Weronika, et al.
Published: (2025)
Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions
by: Sobieski, Bartlomiej, et al.
Published: (2026)
by: Sobieski, Bartlomiej, et al.
Published: (2026)
Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering
by: Chatzoudis, Gerasimos, et al.
Published: (2025)
by: Chatzoudis, Gerasimos, et al.
Published: (2025)
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts
by: Crabbé, Jonathan, et al.
Published: (2023)
by: Crabbé, Jonathan, et al.
Published: (2023)
Aggregated Attributions for Explanatory Analysis of 3D Segmentation Models
by: Chrabaszcz, Maciej, et al.
Published: (2024)
by: Chrabaszcz, Maciej, et al.
Published: (2024)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
by: Pach, Mateusz, et al.
Published: (2025)
by: Pach, Mateusz, et al.
Published: (2025)
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
by: Surkov, Viacheslav, et al.
Published: (2024)
by: Surkov, Viacheslav, et al.
Published: (2024)
SALVE: Sparse Autoencoder-Latent Vector Editing for Mechanistic Control of Neural Networks
by: Flovik, Vegard
Published: (2025)
by: Flovik, Vegard
Published: (2025)
DiffCLIP: Differential Attention Meets CLIP
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
by: Hua, Zhenglin, et al.
Published: (2025)
by: Hua, Zhenglin, et al.
Published: (2025)
Interpretable Prediction of Lymph Node Metastasis in Rectal Cancer MRI Using Variational Autoencoders
by: Keel, Benjamin, et al.
Published: (2025)
by: Keel, Benjamin, et al.
Published: (2025)
CountCLIP -- [Re] Teaching CLIP to Count to Ten
by: Mestha, Harshvardhan, et al.
Published: (2024)
by: Mestha, Harshvardhan, et al.
Published: (2024)
Steering Sparse Autoencoder Latents to Control Dynamic Head Pruning in Vision Transformers (Student Abstract)
by: Lee, Yousung, et al.
Published: (2026)
by: Lee, Yousung, et al.
Published: (2026)
CNN-based explanation ensembling for dataset, representation and explanations evaluation
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
by: Lai, Zhengfeng, et al.
Published: (2023)
by: Lai, Zhengfeng, et al.
Published: (2023)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
GeoSAE: Geometric Prior-Guided Layer-Wise Sparse Autoencoder Annotation of Brain MRI Foundation Models
by: Nerrise, Favour, et al.
Published: (2026)
by: Nerrise, Favour, et al.
Published: (2026)
CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable, and Controllable Text-Guided Face Manipulation
by: Zhou, Chenliang, et al.
Published: (2022)
by: Zhou, Chenliang, et al.
Published: (2022)
CLIP Can Understand Depth
by: Kim, Sohee, et al.
Published: (2024)
by: Kim, Sohee, et al.
Published: (2024)
ECOR: Explainable CLIP for Object Recognition
by: Rasekh, Ali, et al.
Published: (2024)
by: Rasekh, Ali, et al.
Published: (2024)
Causal Interpretation of Sparse Autoencoder Features in Vision
by: Han, Sangyu, et al.
Published: (2025)
by: Han, Sangyu, et al.
Published: (2025)
Improving the Diffusability of Autoencoders
by: Skorokhodov, Ivan, et al.
Published: (2025)
by: Skorokhodov, Ivan, et al.
Published: (2025)
Implicit Inversion turns CLIP into a Decoder
by: D'Orazio, Antonio, et al.
Published: (2025)
by: D'Orazio, Antonio, et al.
Published: (2025)
Steering CLIP's vision transformer with sparse autoencoders
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
IDEA: Image Description Enhanced CLIP-Adapter
by: Ye, Zhipeng, et al.
Published: (2025)
by: Ye, Zhipeng, et al.
Published: (2025)
Similar Items
-
SwordBench: Evaluating Orthogonality of Steering Image Representations
by: Zaigrajew, Vladimir, et al.
Published: (2026) -
Adversarial attacks and defenses in explainable artificial intelligence: A survey
by: Baniecki, Hubert, et al.
Published: (2023) -
Birds look like cars: Adversarial analysis of intrinsically interpretable deep learning
by: Baniecki, Hubert, et al.
Published: (2025) -
LINE: LLM-based Iterative Neuron Explanations for Vision Models
by: Zaigrajew, Vladimir, et al.
Published: (2026) -
Red Teaming Models for Hyperspectral Image Analysis Using Explainable AI
by: Zaigrajew, Vladimir, et al.
Published: (2024)