Beyond Semantics: Disentangling Information Scope in Sparse Autoencoders for CLIP

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ro, Yusung, Choi, Jaehyun, Kim, Junmo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908942613872640
author Ro, Yusung
Choi, Jaehyun
Kim, Junmo
author_facet Ro, Yusung
Choi, Jaehyun
Kim, Junmo
contents Sparse Autoencoders (SAEs) have emerged as a powerful tool for interpreting the internal representations of CLIP vision encoders, yet existing analyses largely focus on the semantic meaning of individual features. We introduce information scope as a complementary dimension of interpretability that characterizes how broadly an SAE feature aggregates visual evidence, ranging from localized, patch-specific cues to global, image-level signals. We observe that some SAE features respond consistently across spatial perturbations, while others shift unpredictably with minor input changes, indicating a fundamental distinction in their underlying scope. To quantify this, we propose the Contextual Dependency Score (CDS), which separates positionally stable local scope features from positionally variant global scope features. Our experiments show that features of different information scopes exert systematically different influences on CLIP's predictions and confidence. These findings establish information scope as a critical new axis for understanding CLIP representations and provide a deeper diagnostic view of SAE-derived features.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05724
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Semantics: Disentangling Information Scope in Sparse Autoencoders for CLIP
Ro, Yusung
Choi, Jaehyun
Kim, Junmo
Computer Vision and Pattern Recognition
Sparse Autoencoders (SAEs) have emerged as a powerful tool for interpreting the internal representations of CLIP vision encoders, yet existing analyses largely focus on the semantic meaning of individual features. We introduce information scope as a complementary dimension of interpretability that characterizes how broadly an SAE feature aggregates visual evidence, ranging from localized, patch-specific cues to global, image-level signals. We observe that some SAE features respond consistently across spatial perturbations, while others shift unpredictably with minor input changes, indicating a fundamental distinction in their underlying scope. To quantify this, we propose the Contextual Dependency Score (CDS), which separates positionally stable local scope features from positionally variant global scope features. Our experiments show that features of different information scopes exert systematically different influences on CLIP's predictions and confidence. These findings establish information scope as a critical new axis for understanding CLIP representations and provide a deeper diagnostic view of SAE-derived features.
title Beyond Semantics: Disentangling Information Scope in Sparse Autoencoders for CLIP
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.05724