From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dreyer, Maximilian, Hufe, Lorenz, Berend, Jim, Wiegand, Thomas, Lapuschkin, Sebastian, Samek, Wojciech
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909623624138752
author Dreyer, Maximilian
Hufe, Lorenz
Berend, Jim
Wiegand, Thomas
Lapuschkin, Sebastian
Samek, Wojciech
author_facet Dreyer, Maximilian
Hufe, Lorenz
Berend, Jim
Wiegand, Thomas
Lapuschkin, Sebastian
Samek, Wojciech
contents Transformer-based CLIP models are widely used for text-image probing and feature extraction, making it relevant to understand the internal mechanisms behind their predictions. While recent works show that Sparse Autoencoders (SAEs) yield interpretable latent components, they focus on what these encode and miss how they drive predictions. We introduce a scalable framework that reveals what latent components activate for, how they align with expected semantics, and how important they are to predictions. To achieve this, we adapt attribution patching for instance-wise component attributions in CLIP and highlight key faithfulness limitations of the widely used Logit Lens technique. By combining attributions with semantic alignment scores, we can automatically uncover reliance on components that encode semantically unexpected or spurious concepts. Applied across multiple CLIP variants, our method uncovers hundreds of surprising components linked to polysemous words, compound nouns, visual typography and dataset artifacts. While text embeddings remain prone to semantic ambiguity, they are more robust to spurious correlations compared to linear classifiers trained on image embeddings. A case study on skin lesion detection highlights how such classifiers can amplify hidden shortcuts, underscoring the need for holistic, mechanistic interpretability. We provide code at https://github.com/maxdreyer/attributing-clip.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20229
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
Dreyer, Maximilian
Hufe, Lorenz
Berend, Jim
Wiegand, Thomas
Lapuschkin, Sebastian
Samek, Wojciech
Machine Learning
Artificial Intelligence
Transformer-based CLIP models are widely used for text-image probing and feature extraction, making it relevant to understand the internal mechanisms behind their predictions. While recent works show that Sparse Autoencoders (SAEs) yield interpretable latent components, they focus on what these encode and miss how they drive predictions. We introduce a scalable framework that reveals what latent components activate for, how they align with expected semantics, and how important they are to predictions. To achieve this, we adapt attribution patching for instance-wise component attributions in CLIP and highlight key faithfulness limitations of the widely used Logit Lens technique. By combining attributions with semantic alignment scores, we can automatically uncover reliance on components that encode semantically unexpected or spurious concepts. Applied across multiple CLIP variants, our method uncovers hundreds of surprising components linked to polysemous words, compound nouns, visual typography and dataset artifacts. While text embeddings remain prone to semantic ambiguity, they are more robust to spurious correlations compared to linear classifiers trained on image embeddings. A case study on skin lesion detection highlights how such classifiers can amplify hidden shortcuts, underscoring the need for holistic, mechanistic interpretability. We provide code at https://github.com/maxdreyer/attributing-clip.
title From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.20229