Processing and acquisition traces in visual encoders: What does CLIP know about your camera?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ramos, Ryan, Stojnić, Vladan, Kordopatis-Zilos, Giorgos, Nakashima, Yuta, Tolias, Giorgos, Garcia, Noa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914436392943616
author Ramos, Ryan
Stojnić, Vladan
Kordopatis-Zilos, Giorgos
Nakashima, Yuta
Tolias, Giorgos
Garcia, Noa
author_facet Ramos, Ryan
Stojnić, Vladan
Kordopatis-Zilos, Giorgos
Nakashima, Yuta
Tolias, Giorgos
Garcia, Noa
contents Prior work has analyzed the robustness of visual encoders to image transformations and corruptions, particularly in cases where such alterations are not seen during training. When this occurs, they introduce a form of distribution shift at test time, often leading to performance degradation. The primary focus has been on severe corruptions that, when applied aggressively, distort useful signals necessary for accurate semantic predictions. We take a different perspective by analyzing parameters of the image acquisition process and transformations that may be subtle or even imperceptible to the human eye. We find that such parameters are systematically encoded in the learned visual representations and can be easily recovered. More strikingly, their presence can have a profound impact, either positively or negatively, on semantic predictions. This effect depends on whether there is a strong correlation or anti-correlation between semantic labels and these acquisition-based or processing-based labels. Our code and data are available at: https://github.com/ryan-caesar-ramos/visual-encoder-traces
format Preprint
id arxiv_https___arxiv_org_abs_2508_10637
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
Ramos, Ryan
Stojnić, Vladan
Kordopatis-Zilos, Giorgos
Nakashima, Yuta
Tolias, Giorgos
Garcia, Noa
Computer Vision and Pattern Recognition
Prior work has analyzed the robustness of visual encoders to image transformations and corruptions, particularly in cases where such alterations are not seen during training. When this occurs, they introduce a form of distribution shift at test time, often leading to performance degradation. The primary focus has been on severe corruptions that, when applied aggressively, distort useful signals necessary for accurate semantic predictions. We take a different perspective by analyzing parameters of the image acquisition process and transformations that may be subtle or even imperceptible to the human eye. We find that such parameters are systematically encoded in the learned visual representations and can be easily recovered. More strikingly, their presence can have a profound impact, either positively or negatively, on semantic predictions. This effect depends on whether there is a strong correlation or anti-correlation between semantic labels and these acquisition-based or processing-based labels. Our code and data are available at: https://github.com/ryan-caesar-ramos/visual-encoder-traces
title Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.10637