Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Grzywaczewski, Jakub, Płudowski, Dawid, Biecek, Przemysław |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SwordBench: Evaluating Orthogonality of Steering Image Representations
by: Zaigrajew, Vladimir, et al.
Published: (2026)
by: Zaigrajew, Vladimir, et al.
Published: (2026)
Rethinking Visual Counterfactual Explanations Through Region Constraint
by: Sobieski, Bartlomiej, et al.
Published: (2024)
by: Sobieski, Bartlomiej, et al.
Published: (2024)
Local Intrinsic Dimension Unveils Hallucinations in Diffusion Models
by: Sobieski, Bartlomiej, et al.
Published: (2026)
by: Sobieski, Bartlomiej, et al.
Published: (2026)
Interpreting CLIP with Hierarchical Sparse Autoencoders
by: Zaigrajew, Vladimir, et al.
Published: (2025)
by: Zaigrajew, Vladimir, et al.
Published: (2025)
Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions
by: Sobieski, Bartlomiej, et al.
Published: (2026)
by: Sobieski, Bartlomiej, et al.
Published: (2026)
Position: Do Not Explain Vision Models Without Context
by: Tomaszewska, Paulina, et al.
Published: (2024)
by: Tomaszewska, Paulina, et al.
Published: (2024)
Global Counterfactual Directions
by: Sobieski, Bartlomiej, et al.
Published: (2024)
by: Sobieski, Bartlomiej, et al.
Published: (2024)
Adversarial attacks and defenses in explainable artificial intelligence: A survey
by: Baniecki, Hubert, et al.
Published: (2023)
by: Baniecki, Hubert, et al.
Published: (2023)
CNN-based explanation ensembling for dataset, representation and explanations evaluation
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
Steering CLIP's vision transformer with sparse autoencoders
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
NormEnsembleXAI: Unveiling the Strengths and Weaknesses of XAI Ensemble Techniques
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
X-ray transferable polyrepresentation learning
by: Hryniewska-Guzik, Weronika, et al.
Published: (2025)
by: Hryniewska-Guzik, Weronika, et al.
Published: (2025)
Red Teaming Models for Hyperspectral Image Analysis Using Explainable AI
by: Zaigrajew, Vladimir, et al.
Published: (2024)
by: Zaigrajew, Vladimir, et al.
Published: (2024)
Birds look like cars: Adversarial analysis of intrinsically interpretable deep learning
by: Baniecki, Hubert, et al.
Published: (2025)
by: Baniecki, Hubert, et al.
Published: (2025)
Self-supervised pretraining for an iterative image size agnostic vision transformer
by: Prisadnikov, Nedyalko, et al.
Published: (2026)
by: Prisadnikov, Nedyalko, et al.
Published: (2026)
Red-Teaming Segment Anything Model
by: Jankowski, Krzysztof, et al.
Published: (2024)
by: Jankowski, Krzysztof, et al.
Published: (2024)
AEGIS: Preserving privacy of 3D Facial Avatars with Adversarial Perturbations
by: Wolkiewicz, Dawid, et al.
Published: (2025)
by: Wolkiewicz, Dawid, et al.
Published: (2025)
Does context matter in digital pathology?
by: Tomaszewska, Paulina, et al.
Published: (2024)
by: Tomaszewska, Paulina, et al.
Published: (2024)
Harnessing small projectors and multiple views for efficient vision pretraining
by: Agrawal, Kumar Krishna, et al.
Published: (2023)
by: Agrawal, Kumar Krishna, et al.
Published: (2023)
LINE: LLM-based Iterative Neuron Explanations for Vision Models
by: Zaigrajew, Vladimir, et al.
Published: (2026)
by: Zaigrajew, Vladimir, et al.
Published: (2026)
The effectiveness of MAE pre-pretraining for billion-scale pretraining
by: Singh, Mannat, et al.
Published: (2023)
by: Singh, Mannat, et al.
Published: (2023)
CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
by: Yang, Shengzhu, et al.
Published: (2025)
by: Yang, Shengzhu, et al.
Published: (2025)
The Dark Patterns of Personalized Persuasion in Large Language Models: Exposing Persuasive Linguistic Features for Big Five Personality Traits in LLMs Responses
by: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Published: (2024)
by: Mieleszczenko-Kowszewicz, Wiktoria, et al.
Published: (2024)
Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions
by: Baniecki, Hubert, et al.
Published: (2025)
by: Baniecki, Hubert, et al.
Published: (2025)
Interpretable machine learning for time-to-event prediction in medicine and healthcare
by: Baniecki, Hubert, et al.
Published: (2023)
by: Baniecki, Hubert, et al.
Published: (2023)
SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
by: Xie, Shaoan, et al.
Published: (2025)
by: Xie, Shaoan, et al.
Published: (2025)
Swin SMT: Global Sequential Modeling in 3D Medical Image Segmentation
by: Płotka, Szymon, et al.
Published: (2024)
by: Płotka, Szymon, et al.
Published: (2024)
A comparative analysis of deep learning models for lung segmentation on X-ray images
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
by: Hryniewska-Guzik, Weronika, et al.
Published: (2024)
Deep spatial context: when attention-based models meet spatial regression
by: Tomaszewska, Paulina, et al.
Published: (2024)
by: Tomaszewska, Paulina, et al.
Published: (2024)
Context-self contrastive pretraining for crop type semantic segmentation
by: Tarasiou, Michail, et al.
Published: (2021)
by: Tarasiou, Michail, et al.
Published: (2021)
UnGuide: Learning to Forget with LoRA-Guided Diffusion Models
by: Polowczyk, Agnieszka, et al.
Published: (2025)
by: Polowczyk, Agnieszka, et al.
Published: (2025)
Are vision language models robust to uncertain inputs?
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
Neural Surface Priors for Editable Gaussian Splatting
by: Szymkowiak, Jakub, et al.
Published: (2024)
by: Szymkowiak, Jakub, et al.
Published: (2024)
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
by: Shi, Danli, et al.
Published: (2024)
by: Shi, Danli, et al.
Published: (2024)
Aggregated Attributions for Explanatory Analysis of 3D Segmentation Models
by: Chrabaszcz, Maciej, et al.
Published: (2024)
by: Chrabaszcz, Maciej, et al.
Published: (2024)
Quantifying the human visual exposome with vision language models
by: Rominger, Christian, et al.
Published: (2026)
by: Rominger, Christian, et al.
Published: (2026)
What matters when building vision-language models?
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
Interpreting vision transformers via residual replacement model
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
by: Garrido, Quentin, et al.
Published: (2025)
by: Garrido, Quentin, et al.
Published: (2025)
Reproducible scaling laws for contrastive language-image learning
by: Cherti, Mehdi, et al.
Published: (2022)
by: Cherti, Mehdi, et al.
Published: (2022)
Similar Items
-
SwordBench: Evaluating Orthogonality of Steering Image Representations
by: Zaigrajew, Vladimir, et al.
Published: (2026) -
Rethinking Visual Counterfactual Explanations Through Region Constraint
by: Sobieski, Bartlomiej, et al.
Published: (2024) -
Local Intrinsic Dimension Unveils Hallucinations in Diffusion Models
by: Sobieski, Bartlomiej, et al.
Published: (2026) -
Interpreting CLIP with Hierarchical Sparse Autoencoders
by: Zaigrajew, Vladimir, et al.
Published: (2025) -
Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions
by: Sobieski, Bartlomiej, et al.
Published: (2026)