CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abbasi, Reza, Nazari, Ali, Sefid, Aminreza, Banayeeanzade, Mohammadali, Rohban, Mohammad Hossein, Baghshah, Mahdieh Soleymani
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909517688602624
author Abbasi, Reza
Nazari, Ali
Sefid, Aminreza
Banayeeanzade, Mohammadali
Rohban, Mohammad Hossein
Baghshah, Mahdieh Soleymani
author_facet Abbasi, Reza
Nazari, Ali
Sefid, Aminreza
Banayeeanzade, Mohammadali
Rohban, Mohammad Hossein
Baghshah, Mahdieh Soleymani
contents Contrastive Language-Image Pre-training (CLIP) models excel in zero-shot classification, yet face challenges in complex multi-object scenarios. This study offers a comprehensive analysis of CLIP's limitations in these contexts using a specialized dataset, ComCO, designed to evaluate CLIP's encoders in diverse multi-object scenarios. Our findings reveal significant biases: the text encoder prioritizes first-mentioned objects, and the image encoder favors larger objects. Through retrieval and classification tasks, we quantify these biases across multiple CLIP variants and trace their origins to CLIP's training process, supported by analyses of the LAION dataset and training progression. Our image-text matching experiments show substantial performance drops when object size or token order changes, underscoring CLIP's instability with rephrased but semantically similar captions. Extending this to longer captions and text-to-image models like Stable Diffusion, we demonstrate how prompt order influences object prominence in generated images. For more details and access to our dataset and analysis code, visit our project repository: https://clip-oscope.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19842
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation
Abbasi, Reza
Nazari, Ali
Sefid, Aminreza
Banayeeanzade, Mohammadali
Rohban, Mohammad Hossein
Baghshah, Mahdieh Soleymani
Computer Vision and Pattern Recognition
Contrastive Language-Image Pre-training (CLIP) models excel in zero-shot classification, yet face challenges in complex multi-object scenarios. This study offers a comprehensive analysis of CLIP's limitations in these contexts using a specialized dataset, ComCO, designed to evaluate CLIP's encoders in diverse multi-object scenarios. Our findings reveal significant biases: the text encoder prioritizes first-mentioned objects, and the image encoder favors larger objects. Through retrieval and classification tasks, we quantify these biases across multiple CLIP variants and trace their origins to CLIP's training process, supported by analyses of the LAION dataset and training progression. Our image-text matching experiments show substantial performance drops when object size or token order changes, underscoring CLIP's instability with rephrased but semantically similar captions. Extending this to longer captions and text-to-image models like Stable Diffusion, we demonstrate how prompt order influences object prominence in generated images. For more details and access to our dataset and analysis code, visit our project repository: https://clip-oscope.github.io.
title CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.19842