Image Reconstruction as a Tool for Feature Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Allakhverdov, Eduard, Tarasov, Dmitrii, Goncharova, Elizaveta, Kuznetsov, Andrey
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918050943467520
author Allakhverdov, Eduard
Tarasov, Dmitrii
Goncharova, Elizaveta
Kuznetsov, Andrey
author_facet Allakhverdov, Eduard
Tarasov, Dmitrii
Goncharova, Elizaveta
Kuznetsov, Andrey
contents Vision encoders are increasingly used in modern applications, from vision-only models to multimodal systems such as vision-language models. Despite their remarkable success, it remains unclear how these architectures represent features internally. Here, we propose a novel approach for interpreting vision features via image reconstruction. We compare two related model families, SigLIP and SigLIP2, which differ only in their training objective, and show that encoders pre-trained on image-based tasks retain significantly more image information than those trained on non-image tasks such as contrastive learning. We further apply our method to a range of vision encoders, ranking them by the informativeness of their feature representations. Finally, we demonstrate that manipulating the feature space yields predictable changes in reconstructed images, revealing that orthogonal rotations (rather than spatial transformations) control color encoding. Our approach can be applied to any vision encoder, shedding light on the inner structure of its feature space. The code and model weights to reproduce the experiments are available in GitHub.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07803
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Image Reconstruction as a Tool for Feature Analysis
Allakhverdov, Eduard
Tarasov, Dmitrii
Goncharova, Elizaveta
Kuznetsov, Andrey
Computer Vision and Pattern Recognition
68T10, 68T30, 68T45
I.2.10
Vision encoders are increasingly used in modern applications, from vision-only models to multimodal systems such as vision-language models. Despite their remarkable success, it remains unclear how these architectures represent features internally. Here, we propose a novel approach for interpreting vision features via image reconstruction. We compare two related model families, SigLIP and SigLIP2, which differ only in their training objective, and show that encoders pre-trained on image-based tasks retain significantly more image information than those trained on non-image tasks such as contrastive learning. We further apply our method to a range of vision encoders, ranking them by the informativeness of their feature representations. Finally, we demonstrate that manipulating the feature space yields predictable changes in reconstructed images, revealing that orthogonal rotations (rather than spatial transformations) control color encoding. Our approach can be applied to any vision encoder, shedding light on the inner structure of its feature space. The code and model weights to reproduce the experiments are available in GitHub.
title Image Reconstruction as a Tool for Feature Analysis
topic Computer Vision and Pattern Recognition
68T10, 68T30, 68T45
I.2.10
url https://arxiv.org/abs/2506.07803