Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chew, Oscar, Huang, Hsiao-Ying, Jain, Kunal, Chen, Tai-I, Doan, Khoa D, Huang, Kuan-Hao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917388429033472
author Chew, Oscar
Huang, Hsiao-Ying
Jain, Kunal
Chen, Tai-I
Doan, Khoa D
Huang, Kuan-Hao
author_facet Chew, Oscar
Huang, Hsiao-Ying
Jain, Kunal
Chen, Tai-I
Doan, Khoa D
Huang, Kuan-Hao
contents Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While a growing body of work has sought to address this limitation, we identify a distinct failure mode in the CLIP family, which we term center bias, that persists even in recent model variants. Specifically, CLIP tends to disproportionately focus on the central region of an image, overlooking important objects located near the boundaries. This limitation is fundamental as failure to recognize relevant objects makes it difficult to perform any sophisticated tasks that depend on those objects. To understand the underlying causes of the limitation, we conduct analyses from both representation and attention perspectives. Using interpretability methods, i.e., embedding decomposition and attention map analysis, we find that relevant concepts especially those associated with off-center objects vanish from the model's embedding in the final representation due to information loss during the aggregation of visual embeddings, particularly the reliance on pooling mechanisms. Finally, we show that this bias can be alleviated with training-free strategies such as visual prompting and attention redistribution by redirecting models' attention to off-center regions.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05971
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family
Chew, Oscar
Huang, Hsiao-Ying
Jain, Kunal
Chen, Tai-I
Doan, Khoa D
Huang, Kuan-Hao
Computer Vision and Pattern Recognition
Computation and Language
Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While a growing body of work has sought to address this limitation, we identify a distinct failure mode in the CLIP family, which we term center bias, that persists even in recent model variants. Specifically, CLIP tends to disproportionately focus on the central region of an image, overlooking important objects located near the boundaries. This limitation is fundamental as failure to recognize relevant objects makes it difficult to perform any sophisticated tasks that depend on those objects. To understand the underlying causes of the limitation, we conduct analyses from both representation and attention perspectives. Using interpretability methods, i.e., embedding decomposition and attention map analysis, we find that relevant concepts especially those associated with off-center objects vanish from the model's embedding in the final representation due to information loss during the aggregation of visual embeddings, particularly the reliance on pooling mechanisms. Finally, we show that this bias can be alleviated with training-free strategies such as visual prompting and attention redistribution by redirecting models' attention to off-center regions.
title Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2604.05971