Embedding Shift Dissection on CLIP: Effects of Augmentations on VLM's Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dahal, Ashim, Murad, Saydul Akbar, Rahimi, Nick
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911532414140416
author Dahal, Ashim
Murad, Saydul Akbar
Rahimi, Nick
author_facet Dahal, Ashim
Murad, Saydul Akbar
Rahimi, Nick
contents Understanding the representation shift on Vision Language Models like CLIP under different augmentations provides valuable insights on Mechanistic Interpretability. In this study, we show the shift on CLIP's embeddings on 9 common augmentation techniques: noise, blur, color jitter, scale and rotate, flip, elastic and perspective transforms, random brightness and contrast, and coarse dropout of pixel blocks. We scrutinize the embedding shifts under similarity on attention map, patch, edge, detail preservation, cosine similarity, L2 distance, pairwise distance and dendrogram clusters and provide qualitative analysis on sample images. Our findings suggest certain augmentations like noise, perspective transform and shift scaling have higher degree of drastic impact on embedding shift. This study provides a concrete foundation for future work on VLM's robustness for mechanical interpretation and adversarial data defense. The code implementation for this study can be found on \href{https://github.com/ashimdahal/clip-shift-analysis}{https://github.com/ashimdahal/clip-shift-analysis}.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23495
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Embedding Shift Dissection on CLIP: Effects of Augmentations on VLM's Representation Learning
Dahal, Ashim
Murad, Saydul Akbar
Rahimi, Nick
Computer Vision and Pattern Recognition
Understanding the representation shift on Vision Language Models like CLIP under different augmentations provides valuable insights on Mechanistic Interpretability. In this study, we show the shift on CLIP's embeddings on 9 common augmentation techniques: noise, blur, color jitter, scale and rotate, flip, elastic and perspective transforms, random brightness and contrast, and coarse dropout of pixel blocks. We scrutinize the embedding shifts under similarity on attention map, patch, edge, detail preservation, cosine similarity, L2 distance, pairwise distance and dendrogram clusters and provide qualitative analysis on sample images. Our findings suggest certain augmentations like noise, perspective transform and shift scaling have higher degree of drastic impact on embedding shift. This study provides a concrete foundation for future work on VLM's robustness for mechanical interpretation and adversarial data defense. The code implementation for this study can be found on \href{https://github.com/ashimdahal/clip-shift-analysis}{https://github.com/ashimdahal/clip-shift-analysis}.
title Embedding Shift Dissection on CLIP: Effects of Augmentations on VLM's Representation Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.23495