CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Abbasi, Reza, Nazari, Ali, Sefid, Aminreza, Banayeeanzade, Mohammadali, Rohban, Mohammad Hossein, Baghshah, Mahdieh Soleymani |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analyzing CLIP's Performance Limitations in Multi-Object Scenarios: A Controlled High-Resolution Study
by: Abbasi, Reza, et al.
Published: (2025)
by: Abbasi, Reza, et al.
Published: (2025)
Deciphering the Role of Representation Disentanglement: Investigating Compositional Generalization in CLIP Models
by: Abbasi, Reza, et al.
Published: (2024)
by: Abbasi, Reza, et al.
Published: (2024)
Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models!
by: Marioriyad, Arash, et al.
Published: (2024)
by: Marioriyad, Arash, et al.
Published: (2024)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
by: Abbasi, Reza, et al.
Published: (2024)
by: Abbasi, Reza, et al.
Published: (2024)
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models
by: Rezaei, Parham, et al.
Published: (2025)
by: Rezaei, Parham, et al.
Published: (2025)
Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models
by: Marioriyad, Arash, et al.
Published: (2024)
by: Marioriyad, Arash, et al.
Published: (2024)
Causal Attribution via Activation Patching
by: Izadi, Amirmohammad, et al.
Published: (2026)
by: Izadi, Amirmohammad, et al.
Published: (2026)
T2I-FineEval: Fine-Grained Compositional Metric for Text-to-Image Evaluation
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2025)
by: Hosseini, Seyed Mohammad Hadi, et al.
Published: (2025)
GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models
by: Abdollahi, Ali, et al.
Published: (2024)
by: Abdollahi, Ali, et al.
Published: (2024)
No Concept Left Behind: Test-Time Optimization for Compositional Text-to-Image Generation
by: Sameti, Mohammad Hossein, et al.
Published: (2025)
by: Sameti, Mohammad Hossein, et al.
Published: (2025)
Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
by: Zohrabi, Reihaneh, et al.
Published: (2025)
by: Zohrabi, Reihaneh, et al.
Published: (2025)
Hidden Meanings in Plain Sight: RebusBench for Evaluating Cognitive Visual Reasoning
by: Kasaei, Seyed Amir, et al.
Published: (2026)
by: Kasaei, Seyed Amir, et al.
Published: (2026)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
by: Kasaei, Seyed Amir, et al.
Published: (2025)
by: Kasaei, Seyed Amir, et al.
Published: (2025)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
by: Izadi, Amirmohammad, et al.
Published: (2025)
by: Izadi, Amirmohammad, et al.
Published: (2025)
Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation
by: Izadi, Amir Mohammad, et al.
Published: (2025)
by: Izadi, Amir Mohammad, et al.
Published: (2025)
RODEO: Robust Outlier Detection via Exposing Adaptive Out-of-Distribution Samples
by: Mirzaei, Hossein, et al.
Published: (2025)
by: Mirzaei, Hossein, et al.
Published: (2025)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
by: Marioriyad, Arash, et al.
Published: (2025)
by: Marioriyad, Arash, et al.
Published: (2025)
Classification of Breast Cancer Histopathology Images using a Modified Supervised Contrastive Learning Method
by: Sani, Matina Mahdizadeh, et al.
Published: (2024)
by: Sani, Matina Mahdizadeh, et al.
Published: (2024)
CARINOX: Inference-time Scaling with Category-Aware Reward-based Initial Noise Optimization and Exploration
by: Kasaei, Seyed Amir, et al.
Published: (2025)
by: Kasaei, Seyed Amir, et al.
Published: (2025)
Infinity and Beyond: Compositional Alignment in VAR and Diffusion T2I Models
by: Shahabadi, Hossein, et al.
Published: (2025)
by: Shahabadi, Hossein, et al.
Published: (2025)
LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions
by: Mehri, Faridoun, et al.
Published: (2024)
by: Mehri, Faridoun, et al.
Published: (2024)
Dilated Balanced Cross Entropy Loss for Medical Image Segmentation
by: Hosseini, Seyed Mohsen, et al.
Published: (2024)
by: Hosseini, Seyed Mohsen, et al.
Published: (2024)
Improving 3D Few-Shot Segmentation with Inference-Time Pseudo-Labeling
by: Mozafari, Mohammad, et al.
Published: (2024)
by: Mozafari, Mohammad, et al.
Published: (2024)
Lying to Win: Assessing LLM Deception through Human-AI Games and Parallel-World Probing
by: Marioriyad, Arash, et al.
Published: (2026)
by: Marioriyad, Arash, et al.
Published: (2026)
Trained Models Tell Us How to Make Them Robust to Spurious Correlation without Group Annotation
by: Ghaznavi, Mahdi, et al.
Published: (2024)
by: Ghaznavi, Mahdi, et al.
Published: (2024)
ComAlign: Compositional Alignment in Vision-Language Models
by: Abdollah, Ali, et al.
Published: (2024)
by: Abdollah, Ali, et al.
Published: (2024)
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
by: Zohrabi, Reihaneh, et al.
Published: (2026)
by: Zohrabi, Reihaneh, et al.
Published: (2026)
A Contrastive Teacher-Student Framework for Novelty Detection under Style Shifts
by: Mirzaei, Hossein, et al.
Published: (2025)
by: Mirzaei, Hossein, et al.
Published: (2025)
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
by: Marioriyad, Arash, et al.
Published: (2026)
by: Marioriyad, Arash, et al.
Published: (2026)
Understanding Counting Mechanisms in Large Language and Vision-Language Models
by: Hasani, Hosein, et al.
Published: (2025)
by: Hasani, Hosein, et al.
Published: (2025)
Uncovering Grounding IDs: How External Cues Shape Multimodal Binding
by: Hasani, Hosein, et al.
Published: (2025)
by: Hasani, Hosein, et al.
Published: (2025)
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
by: Monsefi, Amin Karimi, et al.
Published: (2024)
by: Monsefi, Amin Karimi, et al.
Published: (2024)
Novel Pipeline for Diagnosing Acute Lymphoblastic Leukemia Sensitive to Related Biomarkers
by: Farsangi, Amirhossein Askari, et al.
Published: (2023)
by: Farsangi, Amirhossein Askari, et al.
Published: (2023)
Hallucination as an Upper Bound: A New Perspective on Text-to-Image Evaluation
by: Kasaei, Seyed Amir, et al.
Published: (2025)
by: Kasaei, Seyed Amir, et al.
Published: (2025)
Decompose-and-Compose: A Compositional Approach to Mitigating Spurious Correlation
by: Noohdani, Fahimeh Hosseini, et al.
Published: (2024)
by: Noohdani, Fahimeh Hosseini, et al.
Published: (2024)
Mitigating Spurious Negative Pairs for Robust Industrial Anomaly Detection
by: Mirzaei, Hossein, et al.
Published: (2025)
by: Mirzaei, Hossein, et al.
Published: (2025)
Large Language Models for Scientific Idea Generation: A Creativity-Centered Survey
by: Shahhosseini, Fatemeh, et al.
Published: (2025)
by: Shahhosseini, Fatemeh, et al.
Published: (2025)
Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy
by: Hasani, Hosein, et al.
Published: (2026)
by: Hasani, Hosein, et al.
Published: (2026)
microCLIP: Unsupervised CLIP Adaptation via Coarse-Fine Token Fusion for Fine-Grained Image Classification
by: Silva, Sathira, et al.
Published: (2025)
by: Silva, Sathira, et al.
Published: (2025)
Patch-Wise Self-Supervised Visual Representation Learning: A Fine-Grained Approach
by: Javidani, Ali, et al.
Published: (2023)
by: Javidani, Ali, et al.
Published: (2023)
Similar Items
-
Analyzing CLIP's Performance Limitations in Multi-Object Scenarios: A Controlled High-Resolution Study
by: Abbasi, Reza, et al.
Published: (2025) -
Deciphering the Role of Representation Disentanglement: Investigating Compositional Generalization in CLIP Models
by: Abbasi, Reza, et al.
Published: (2024) -
Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models!
by: Marioriyad, Arash, et al.
Published: (2024) -
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
by: Abbasi, Reza, et al.
Published: (2024) -
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models
by: Rezaei, Parham, et al.
Published: (2025)