Explaining CLIP Zero-shot Predictions Through Concepts
Fuente:
arXiv
Saved in:
| Main Authors: | Ozdemir, Onat, Christensen, Anders, Alaniz, Stephan, Akata, Zeynep, Akbas, Emre |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Concept Bottleneck Models: From Pitfalls to Solutions
by: Tapli, Merve, et al.
Published: (2026)
by: Tapli, Merve, et al.
Published: (2026)
A Systematic Comparison of Training Objectives for Out-of-Distribution Detection in Image Classification
by: Genç, Furkan, et al.
Published: (2026)
by: Genç, Furkan, et al.
Published: (2026)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
by: Kim, Jae Myung, et al.
Published: (2025)
by: Kim, Jae Myung, et al.
Published: (2025)
DataDream: Few-shot Guided Dataset Generation
by: Kim, Jae Myung, et al.
Published: (2024)
by: Kim, Jae Myung, et al.
Published: (2024)
A Large Scale Analysis of Gender Biases in Text-to-Image Generative Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
by: Kim, Sanghwan, et al.
Published: (2025)
by: Kim, Sanghwan, et al.
Published: (2025)
Concept-Guided Interpretability via Neural Chunking
by: Wu, Shuchen, et al.
Published: (2025)
by: Wu, Shuchen, et al.
Published: (2025)
SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
by: Bader, Jessica, et al.
Published: (2025)
by: Bader, Jessica, et al.
Published: (2025)
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
by: Xiao, Rui, et al.
Published: (2026)
by: Xiao, Rui, et al.
Published: (2026)
FLAIR: VLM with Fine-grained Language-informed Image Representations
by: Xiao, Rui, et al.
Published: (2024)
by: Xiao, Rui, et al.
Published: (2024)
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
by: Kim, Sanghwan, et al.
Published: (2024)
by: Kim, Sanghwan, et al.
Published: (2024)
Unsupervised Image Classification with Adaptive Nearest Neighbor Selection and Cluster Ensembles
by: Baydar, Melih, et al.
Published: (2025)
by: Baydar, Melih, et al.
Published: (2025)
Intrinsic Dimensionality as a Model-Free Measure of Class Imbalance
by: Eser, Çağrı, et al.
Published: (2025)
by: Eser, Çağrı, et al.
Published: (2025)
GenCLIP: Generalizing CLIP Prompts for Zero-shot Anomaly Detection
by: Kim, Donghyeong, et al.
Published: (2025)
by: Kim, Donghyeong, et al.
Published: (2025)
Sparse Autoencoders are Topic Models
by: Girrbach, Leander, et al.
Published: (2025)
by: Girrbach, Leander, et al.
Published: (2025)
Early-exit Convolutional Neural Networks
by: Demir, Edanur, et al.
Published: (2024)
by: Demir, Edanur, et al.
Published: (2024)
CLIP-driven Zero-shot Learning with Ambiguous Labels
by: Fan, Jinfu, et al.
Published: (2026)
by: Fan, Jinfu, et al.
Published: (2026)
SPECIAL: Zero-shot Hyperspectral Image Classification With CLIP
by: Pang, Li, et al.
Published: (2025)
by: Pang, Li, et al.
Published: (2025)
Audio-Visual Generalized Zero-Shot Learning using Pre-Trained Large Multi-Modal Models
by: Kurzendörfer, David, et al.
Published: (2024)
by: Kurzendörfer, David, et al.
Published: (2024)
Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
by: Singhi, Nishad, et al.
Published: (2024)
by: Singhi, Nishad, et al.
Published: (2024)
MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition
by: Song, Dan, et al.
Published: (2023)
by: Song, Dan, et al.
Published: (2023)
MatchED: Crisp Edge Detection Using End-to-End, Matching-based Supervision
by: Cetinkaya, Bedrettin, et al.
Published: (2026)
by: Cetinkaya, Bedrettin, et al.
Published: (2026)
RankED: Addressing Imbalance and Uncertainty in Edge Detection Using Ranking-based Losses
by: Cetinkaya, Bedrettin, et al.
Published: (2024)
by: Cetinkaya, Bedrettin, et al.
Published: (2024)
MIRAGE: Benchmarking and Aligning Multi-Instance Image Editing
by: Liu, Ziqian, et al.
Published: (2026)
by: Liu, Ziqian, et al.
Published: (2026)
Partial CLIP is Enough: Chimera-Seg for Zero-shot Semantic Segmentation
by: Chen, Jialei, et al.
Published: (2025)
by: Chen, Jialei, et al.
Published: (2025)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
by: Xing, Songlong, et al.
Published: (2025)
by: Xing, Songlong, et al.
Published: (2025)
AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection
by: Zhou, Qihang, et al.
Published: (2023)
by: Zhou, Qihang, et al.
Published: (2023)
TokenCLIP: Token-wise Prompt Learning for Zero-shot Anomaly Detection
by: Zhou, Qihang, et al.
Published: (2025)
by: Zhou, Qihang, et al.
Published: (2025)
Unseen No More: Unlocking the Potential of CLIP for Generative Zero-shot HOI Detection
by: Guo, Yixin, et al.
Published: (2024)
by: Guo, Yixin, et al.
Published: (2024)
Interpretable Zero-shot Learning with Infinite Class Concepts
by: Ye, Zihan, et al.
Published: (2025)
by: Ye, Zihan, et al.
Published: (2025)
Vision-by-Language for Training-Free Compositional Image Retrieval
by: Karthik, Shyamgopal, et al.
Published: (2023)
by: Karthik, Shyamgopal, et al.
Published: (2023)
CardiacCLIP: Video-based CLIP Adaptation for LVEF Prediction in a Few-shot Manner
by: Du, Yao, et al.
Published: (2025)
by: Du, Yao, et al.
Published: (2025)
AA-CLIP: Enhancing Zero-shot Anomaly Detection via Anomaly-Aware CLIP
by: Ma, Wenxin, et al.
Published: (2025)
by: Ma, Wenxin, et al.
Published: (2025)
Representation Recycling for Streaming Video Analysis
by: Ertenli, Can Ufuk, et al.
Published: (2022)
by: Ertenli, Can Ufuk, et al.
Published: (2022)
MoCap-to-Visual Domain Adaptation for Efficient Human Mesh Estimation from 2D Keypoints
by: Uguz, Bedirhan, et al.
Published: (2024)
by: Uguz, Bedirhan, et al.
Published: (2024)
Geometry Fidelity for Spherical Images
by: Christensen, Anders, et al.
Published: (2024)
by: Christensen, Anders, et al.
Published: (2024)
Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP
by: Yu, Yating, et al.
Published: (2024)
by: Yu, Yating, et al.
Published: (2024)
DeLoRA: Decoupling Angles and Strength in Low-rank Adaptation
by: Bini, Massimo, et al.
Published: (2025)
by: Bini, Massimo, et al.
Published: (2025)
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
by: Hummel, Thomas, et al.
Published: (2024)
by: Hummel, Thomas, et al.
Published: (2024)
Similar Items
-
Rethinking Concept Bottleneck Models: From Pitfalls to Solutions
by: Tapli, Merve, et al.
Published: (2026) -
A Systematic Comparison of Training Objectives for Out-of-Distribution Detection in Image Classification
by: Genç, Furkan, et al.
Published: (2026) -
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
by: Kim, Jae Myung, et al.
Published: (2025) -
DataDream: Few-shot Guided Dataset Generation
by: Kim, Jae Myung, et al.
Published: (2024) -
A Large Scale Analysis of Gender Biases in Text-to-Image Generative Models
by: Girrbach, Leander, et al.
Published: (2025)