Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Pach, Mateusz, Karthik, Shyamgopal, Bouniot, Quentin, Belongie, Serge, Akata, Zeynep |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
by: Pach, Mateusz, et al.
Published: (2026)
by: Pach, Mateusz, et al.
Published: (2026)
Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
by: Bader, Jessica, et al.
Published: (2025)
by: Bader, Jessica, et al.
Published: (2025)
Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers
by: Roschmann, Simon, et al.
Published: (2025)
by: Roschmann, Simon, et al.
Published: (2025)
Concept-Guided Interpretability via Neural Chunking
by: Wu, Shuchen, et al.
Published: (2025)
by: Wu, Shuchen, et al.
Published: (2025)
Sparse Autoencoders are Topic Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
Noise Hypernetworks: Amortizing Test-Time Compute in Diffusion Models
by: Eyring, Luca, et al.
Published: (2025)
by: Eyring, Luca, et al.
Published: (2025)
Post-hoc Probabilistic Vision-Language Models
by: Baumann, Anton, et al.
Published: (2024)
by: Baumann, Anton, et al.
Published: (2024)
SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport
by: Roschmann, Simon, et al.
Published: (2026)
by: Roschmann, Simon, et al.
Published: (2026)
Vision-by-Language for Training-Free Compositional Image Retrieval
by: Karthik, Shyamgopal, et al.
Published: (2023)
by: Karthik, Shyamgopal, et al.
Published: (2023)
TORE: Token Recycling in Vision Transformers for Efficient Active Visual Exploration
by: Olszewski, Jan, et al.
Published: (2023)
by: Olszewski, Jan, et al.
Published: (2023)
LucidPPN: Unambiguous Prototypical Parts Network for User-centric Interpretable Computer Vision
by: Pach, Mateusz, et al.
Published: (2024)
by: Pach, Mateusz, et al.
Published: (2024)
From Drop-off to Recovery: A Mechanistic Analysis of Segmentation in MLLMs
by: Wu, Boyong, et al.
Published: (2026)
by: Wu, Boyong, et al.
Published: (2026)
Unlearning-based Neural Interpretations
by: Choi, Ching Lam, et al.
Published: (2024)
by: Choi, Ching Lam, et al.
Published: (2024)
Deep Dreams Are Made of This: Visualizing Monosemantic Features in Diffusion Models
by: Szokalski, Adam, et al.
Published: (2026)
by: Szokalski, Adam, et al.
Published: (2026)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
by: Kim, Sanghwan, et al.
Published: (2024)
by: Kim, Sanghwan, et al.
Published: (2024)
Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
by: Singhi, Nishad, et al.
Published: (2024)
by: Singhi, Nishad, et al.
Published: (2024)
Restyling Unsupervised Concept Based Interpretable Networks with Generative Models
by: Parekh, Jayneel, et al.
Published: (2024)
by: Parekh, Jayneel, et al.
Published: (2024)
SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
by: Bader, Jessica, et al.
Published: (2025)
by: Bader, Jessica, et al.
Published: (2025)
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
by: Guimard, Quentin, et al.
Published: (2026)
by: Guimard, Quentin, et al.
Published: (2026)
MMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation Learning
by: Nedungadi, Vishal, et al.
Published: (2024)
by: Nedungadi, Vishal, et al.
Published: (2024)
Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models
by: Yeung, Calvin, et al.
Published: (2026)
by: Yeung, Calvin, et al.
Published: (2026)
Feature Expansion and enhanced Compression for Class Incremental Learning
by: Ferdinand, Quentin, et al.
Published: (2024)
by: Ferdinand, Quentin, et al.
Published: (2024)
Steering Sparse Autoencoder Latents to Control Dynamic Head Pruning in Vision Transformers (Student Abstract)
by: Lee, Yousung, et al.
Published: (2026)
by: Lee, Yousung, et al.
Published: (2026)
Interpreting CLIP with Hierarchical Sparse Autoencoders
by: Zaigrajew, Vladimir, et al.
Published: (2025)
by: Zaigrajew, Vladimir, et al.
Published: (2025)
Unbalancedness in Neural Monge Maps Improves Unpaired Domain Translation
by: Eyring, Luca, et al.
Published: (2023)
by: Eyring, Luca, et al.
Published: (2023)
Simplifying Knowledge Transfer in Pretrained Models
by: Jain, Siddharth, et al.
Published: (2025)
by: Jain, Siddharth, et al.
Published: (2025)
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
by: Surkov, Viacheslav, et al.
Published: (2024)
by: Surkov, Viacheslav, et al.
Published: (2024)
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
by: Hummel, Thomas, et al.
Published: (2024)
by: Hummel, Thomas, et al.
Published: (2024)
Neural Sentinel: Unified Vision Language Model (VLM) for License Plate Recognition with Human-in-the-Loop Continual Learning
by: Sivakoti, Karthik
Published: (2026)
by: Sivakoti, Karthik
Published: (2026)
Rethinking Concept Bottleneck Models: From Pitfalls to Solutions
by: Tapli, Merve, et al.
Published: (2026)
by: Tapli, Merve, et al.
Published: (2026)
ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization
by: Eyring, Luca, et al.
Published: (2024)
by: Eyring, Luca, et al.
Published: (2024)
Stitched Value Model for Diffusion Alignment
by: Go, Hyojun, et al.
Published: (2026)
by: Go, Hyojun, et al.
Published: (2026)
UNBOX: Unveiling Black-box visual models with Natural-language
by: Carnemolla, Simone, et al.
Published: (2026)
by: Carnemolla, Simone, et al.
Published: (2026)
MirrorCheck: Efficient Adversarial Defense for Vision-Language Models
by: Fares, Samar, et al.
Published: (2024)
by: Fares, Samar, et al.
Published: (2024)
Causal Interpretation of Sparse Autoencoder Features in Vision
by: Han, Sangyu, et al.
Published: (2025)
by: Han, Sangyu, et al.
Published: (2025)
SALVE: Sparse Autoencoder-Latent Vector Editing for Mechanistic Control of Neural Networks
by: Flovik, Vegard
Published: (2025)
by: Flovik, Vegard
Published: (2025)
Spurious Feature Eraser: Stabilizing Test-Time Adaptation for Vision-Language Foundation Model
by: Ma, Huan, et al.
Published: (2024)
by: Ma, Huan, et al.
Published: (2024)
GeoSAE: Geometric Prior-Guided Layer-Wise Sparse Autoencoder Annotation of Brain MRI Foundation Models
by: Nerrise, Favour, et al.
Published: (2026)
by: Nerrise, Favour, et al.
Published: (2026)
Compositional Entailment Learning for Hyperbolic Vision-Language Models
by: Pal, Avik, et al.
Published: (2024)
by: Pal, Avik, et al.
Published: (2024)
Tree of Attributes Prompt Learning for Vision-Language Models
by: Ding, Tong, et al.
Published: (2024)
by: Ding, Tong, et al.
Published: (2024)
Similar Items
-
The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
by: Pach, Mateusz, et al.
Published: (2026) -
Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
by: Bader, Jessica, et al.
Published: (2025) -
Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers
by: Roschmann, Simon, et al.
Published: (2025) -
Concept-Guided Interpretability via Neural Chunking
by: Wu, Shuchen, et al.
Published: (2025) -
Sparse Autoencoders are Topic Models
by: Girrbach, Leander, et al.
Published: (2025)