Steering CLIP's vision transformer with sparse autoencoders
Fuente:
arXiv
Guardado en:
| Autores principales: | Joseph, Sonia, Suresh, Praneet, Goldfarb, Ethan, Hufe, Lorenz, Gandelsman, Yossi, Graham, Robert, Bzdok, Danilo, Samek, Wojciech, Richards, Blake Aaron |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
por: Joseph, Sonia, et al.
Publicado: (2025)
por: Joseph, Sonia, et al.
Publicado: (2025)
From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
por: Suresh, Praneet, et al.
Publicado: (2025)
por: Suresh, Praneet, et al.
Publicado: (2025)
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
por: Hufe, Lorenz, et al.
Publicado: (2025)
por: Hufe, Lorenz, et al.
Publicado: (2025)
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
por: Dreyer, Maximilian, et al.
Publicado: (2025)
por: Dreyer, Maximilian, et al.
Publicado: (2025)
Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
por: Bali, Karan, et al.
Publicado: (2026)
por: Bali, Karan, et al.
Publicado: (2026)
Interpreting ResNet-based CLIP via Neuron-Attention Decomposition
por: Bu, Edmund, et al.
Publicado: (2025)
por: Bu, Edmund, et al.
Publicado: (2025)
The Unreasonable Effectiveness of Text Embedding Interpolation for Continuous Image Steering
por: Ekin, Yigit, et al.
Publicado: (2026)
por: Ekin, Yigit, et al.
Publicado: (2026)
Interpreting the Second-Order Effects of Neurons in CLIP
por: Gandelsman, Yossi, et al.
Publicado: (2024)
por: Gandelsman, Yossi, et al.
Publicado: (2024)
Interpreting CLIP's Image Representation via Text-Based Decomposition
por: Gandelsman, Yossi, et al.
Publicado: (2023)
por: Gandelsman, Yossi, et al.
Publicado: (2023)
Quantifying and Enabling the Interpretability of CLIP-like Models
por: Madasu, Avinash, et al.
Publicado: (2024)
por: Madasu, Avinash, et al.
Publicado: (2024)
The Uncanny Valley: A Comprehensive Analysis of Diffusion Models
por: Ghanem, Karam, et al.
Publicado: (2024)
por: Ghanem, Karam, et al.
Publicado: (2024)
Estimating Unknown Population Sizes Using the Hypergeometric Distribution
por: Hodgson, Liam, et al.
Publicado: (2024)
por: Hodgson, Liam, et al.
Publicado: (2024)
Learning Video Representations without Natural Videos
por: Yu, Xueyang, et al.
Publicado: (2024)
por: Yu, Xueyang, et al.
Publicado: (2024)
From Attribution to Action: A Human-Centered Application of Activation Steering
por: Labarta, Tobias, et al.
Publicado: (2026)
por: Labarta, Tobias, et al.
Publicado: (2026)
Vision Transformers Don't Need Trained Registers
por: Jiang, Nick, et al.
Publicado: (2025)
por: Jiang, Nick, et al.
Publicado: (2025)
Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
por: Jiang, Nick, et al.
Publicado: (2024)
por: Jiang, Nick, et al.
Publicado: (2024)
In-Context Representation Hijacking
por: Yona, Itay, et al.
Publicado: (2025)
por: Yona, Itay, et al.
Publicado: (2025)
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
por: Nikankin, Yaniv, et al.
Publicado: (2025)
por: Nikankin, Yaniv, et al.
Publicado: (2025)
Scaling and evaluating sparse autoencoders
por: Gao, Leo, et al.
Publicado: (2024)
por: Gao, Leo, et al.
Publicado: (2024)
The Case for Model Science: Verify, Explore, Steer, Refine
por: Biecek, Przemyslaw, et al.
Publicado: (2026)
por: Biecek, Przemyslaw, et al.
Publicado: (2026)
Model Science: getting serious about verification, explanation and control of AI systems
por: Biecek, Przemyslaw, et al.
Publicado: (2025)
por: Biecek, Przemyslaw, et al.
Publicado: (2025)
Position: Explain to Question not to Justify
por: Biecek, Przemyslaw, et al.
Publicado: (2024)
por: Biecek, Przemyslaw, et al.
Publicado: (2024)
From Weights to Activations: Is Steering the Next Frontier of Adaptation?
por: Ostermann, Simon, et al.
Publicado: (2026)
por: Ostermann, Simon, et al.
Publicado: (2026)
The More You See in 2D, the More You Perceive in 3D
por: Han, Xinyang, et al.
Publicado: (2024)
por: Han, Xinyang, et al.
Publicado: (2024)
LLMs can see and hear without any training
por: Ashutosh, Kumar, et al.
Publicado: (2025)
por: Ashutosh, Kumar, et al.
Publicado: (2025)
Interpreting the Repeated Token Phenomenon in Large Language Models
por: Yona, Itay, et al.
Publicado: (2025)
por: Yona, Itay, et al.
Publicado: (2025)
Jailbreaking Vision-Language Models Through the Visual Modality
por: Azulay, Aharon, et al.
Publicado: (2026)
por: Azulay, Aharon, et al.
Publicado: (2026)
Decomposing multimodal embedding spaces with group-sparse autoencoders
por: Kaushik, Chiraag, et al.
Publicado: (2026)
por: Kaushik, Chiraag, et al.
Publicado: (2026)
Understanding sparse autoencoder scaling in the presence of feature manifolds
por: Michaud, Eric J., et al.
Publicado: (2025)
por: Michaud, Eric J., et al.
Publicado: (2025)
Applying sparse autoencoders to unlearn knowledge in language models
por: Farrell, Eoin, et al.
Publicado: (2024)
por: Farrell, Eoin, et al.
Publicado: (2024)
Iterative Inference in a Chess-Playing Neural Network
por: Sandmann, Elias, et al.
Publicado: (2025)
por: Sandmann, Elias, et al.
Publicado: (2025)
Benchmarking Uncertainty and its Disentanglement in multi-label Chest X-Ray Classification
por: Baur, Simon, et al.
Publicado: (2025)
por: Baur, Simon, et al.
Publicado: (2025)
Learning biologically relevant features in a pathology foundation model using sparse autoencoders
por: Le, Nhat Minh, et al.
Publicado: (2024)
por: Le, Nhat Minh, et al.
Publicado: (2024)
Designing a system of labor market statistics and information / Robert S. Goldfarb, Arvil V. Adams
por: Goldfarb, Robert S
Publicado: (1993)
por: Goldfarb, Robert S
Publicado: (1993)
Towards the AI Historian: Agentic Information Extraction from Primary Sources
por: Hufe, Lorenz, et al.
Publicado: (2026)
por: Hufe, Lorenz, et al.
Publicado: (2026)
Mitochondria‐nucleus crosstalk characterizes Alzheimer's disease across 1,5 million brain cells
por: Chloé Savignac, et al.
Publicado: (2025)
por: Chloé Savignac, et al.
Publicado: (2025)
Can sparse autoencoders be used to decompose and interpret steering vectors?
por: Mayne, Harry, et al.
Publicado: (2024)
por: Mayne, Harry, et al.
Publicado: (2024)
Investigating task-specific prompts and sparse autoencoders for activation monitoring
por: Tillman, Henk, et al.
Publicado: (2025)
por: Tillman, Henk, et al.
Publicado: (2025)
Teaching Humans Subtle Differences with DIFFusion
por: Chiquier, Mia, et al.
Publicado: (2025)
por: Chiquier, Mia, et al.
Publicado: (2025)
An Empirical Study of Autoregressive Pre-training from Videos
por: Rajasegaran, Jathushan, et al.
Publicado: (2025)
por: Rajasegaran, Jathushan, et al.
Publicado: (2025)
Ejemplares similares
-
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
por: Joseph, Sonia, et al.
Publicado: (2025) -
From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
por: Suresh, Praneet, et al.
Publicado: (2025) -
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
por: Hufe, Lorenz, et al.
Publicado: (2025) -
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
por: Dreyer, Maximilian, et al.
Publicado: (2025) -
Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
por: Bali, Karan, et al.
Publicado: (2026)