Steering CLIP's vision transformer with sparse autoencoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Joseph, Sonia, Suresh, Praneet, Goldfarb, Ethan, Hufe, Lorenz, Gandelsman, Yossi, Graham, Robert, Bzdok, Danilo, Samek, Wojciech, Richards, Blake Aaron
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910909308338176
author Joseph, Sonia
Suresh, Praneet
Goldfarb, Ethan
Hufe, Lorenz
Gandelsman, Yossi
Graham, Robert
Bzdok, Danilo
Samek, Wojciech
Richards, Blake Aaron
author_facet Joseph, Sonia
Suresh, Praneet
Goldfarb, Ethan
Hufe, Lorenz
Gandelsman, Yossi
Graham, Robert
Bzdok, Danilo
Samek, Wojciech
Richards, Blake Aaron
contents While vision models are highly capable, their internal mechanisms remain poorly understood -- a challenge which sparse autoencoders (SAEs) have helped address in language, but which remains underexplored in vision. We address this gap by training SAEs on CLIP's vision transformer and uncover key differences between vision and language processing, including distinct sparsity patterns for SAEs trained across layers and token types. We then provide the first systematic analysis on the steerability of CLIP's vision transformer by introducing metrics to quantify how precisely SAE features can be steered to affect the model's output. We find that 10-15\% of neurons and features are steerable, with SAEs providing thousands more steerable features than the base model. Through targeted suppression of SAE features, we then demonstrate improved performance on three vision disentanglement tasks (CelebA, Waterbirds, and typographic attacks), finding optimal disentanglement in middle model layers, and achieving state-of-the-art performance on defense against typographic attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2504_08729
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Steering CLIP's vision transformer with sparse autoencoders
Joseph, Sonia
Suresh, Praneet
Goldfarb, Ethan
Hufe, Lorenz
Gandelsman, Yossi
Graham, Robert
Bzdok, Danilo
Samek, Wojciech
Richards, Blake Aaron
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
While vision models are highly capable, their internal mechanisms remain poorly understood -- a challenge which sparse autoencoders (SAEs) have helped address in language, but which remains underexplored in vision. We address this gap by training SAEs on CLIP's vision transformer and uncover key differences between vision and language processing, including distinct sparsity patterns for SAEs trained across layers and token types. We then provide the first systematic analysis on the steerability of CLIP's vision transformer by introducing metrics to quantify how precisely SAE features can be steered to affect the model's output. We find that 10-15\% of neurons and features are steerable, with SAEs providing thousands more steerable features than the base model. Through targeted suppression of SAE features, we then demonstrate improved performance on three vision disentanglement tasks (CelebA, Waterbirds, and typographic attacks), finding optimal disentanglement in middle model layers, and achieving state-of-the-art performance on defense against typographic attacks.
title Steering CLIP's vision transformer with sparse autoencoders
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2504.08729