Quantifying Structure in CLIP Embeddings: A Statistical Framework for Concept Interpretation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Jitian, Li, Chenghui, Sala, Frederic, Rohe, Karl
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915346980536320
author Zhao, Jitian
Li, Chenghui
Sala, Frederic
Rohe, Karl
author_facet Zhao, Jitian
Li, Chenghui
Sala, Frederic
Rohe, Karl
contents Concept-based approaches, which aim to identify human-understandable concepts within a model's internal representations, are a promising method for interpreting embeddings from deep neural network models, such as CLIP. While these approaches help explain model behavior, current methods lack statistical rigor, making it challenging to validate identified concepts and compare different techniques. To address this challenge, we introduce a hypothesis testing framework that quantifies rotation-sensitive structures within the CLIP embedding space. Once such structures are identified, we propose a post-hoc concept decomposition method. Unlike existing approaches, it offers theoretical guarantees that discovered concepts represent robust, reproducible patterns (rather than method-specific artifacts) and outperforms other techniques in terms of reconstruction error. Empirically, we demonstrate that our concept-based decomposition algorithm effectively balances reconstruction accuracy with concept interpretability and helps mitigate spurious cues in data. Applied to a popular spurious correlation dataset, our method yields a 22.6% increase in worst-group accuracy after removing spurious background concepts.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13831
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Quantifying Structure in CLIP Embeddings: A Statistical Framework for Concept Interpretation
Zhao, Jitian
Li, Chenghui
Sala, Frederic
Rohe, Karl
Machine Learning
Artificial Intelligence
Concept-based approaches, which aim to identify human-understandable concepts within a model's internal representations, are a promising method for interpreting embeddings from deep neural network models, such as CLIP. While these approaches help explain model behavior, current methods lack statistical rigor, making it challenging to validate identified concepts and compare different techniques. To address this challenge, we introduce a hypothesis testing framework that quantifies rotation-sensitive structures within the CLIP embedding space. Once such structures are identified, we propose a post-hoc concept decomposition method. Unlike existing approaches, it offers theoretical guarantees that discovered concepts represent robust, reproducible patterns (rather than method-specific artifacts) and outperforms other techniques in terms of reconstruction error. Empirically, we demonstrate that our concept-based decomposition algorithm effectively balances reconstruction accuracy with concept interpretability and helps mitigate spurious cues in data. Applied to a popular spurious correlation dataset, our method yields a 22.6% increase in worst-group accuracy after removing spurious background concepts.
title Quantifying Structure in CLIP Embeddings: A Statistical Framework for Concept Interpretation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.13831