Refining CLIP's Spatial Awareness: A Visual-Centric Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiu, Congpei, Wu, Yanhao, Ke, Wei, Bai, Xiuxiu, Zhang, Tong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910902585917440
author Qiu, Congpei
Wu, Yanhao
Ke, Wei
Bai, Xiuxiu
Zhang, Tong
author_facet Qiu, Congpei
Wu, Yanhao
Ke, Wei
Bai, Xiuxiu
Zhang, Tong
contents Contrastive Language-Image Pre-training (CLIP) excels in global alignment with language but exhibits limited sensitivity to spatial information, leading to strong performance in zero-shot classification tasks but underperformance in tasks requiring precise spatial understanding. Recent approaches have introduced Region-Language Alignment (RLA) to enhance CLIP's performance in dense multimodal tasks by aligning regional visual representations with corresponding text inputs. However, we find that CLIP ViTs fine-tuned with RLA suffer from notable loss in spatial awareness, which is crucial for dense prediction tasks. To address this, we propose the Spatial Correlation Distillation (SCD) framework, which preserves CLIP's inherent spatial structure and mitigates the above degradation. To further enhance spatial correlations, we introduce a lightweight Refiner that extracts refined correlations directly from CLIP before feeding them into SCD, based on an intriguing finding that CLIP naturally captures high-quality dense features. Together, these components form a robust distillation framework that enables CLIP ViTs to integrate both visual-language and visual-centric improvements, achieving state-of-the-art results across various open-vocabulary dense prediction benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2504_02328
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Refining CLIP's Spatial Awareness: A Visual-Centric Perspective
Qiu, Congpei
Wu, Yanhao
Ke, Wei
Bai, Xiuxiu
Zhang, Tong
Computer Vision and Pattern Recognition
Contrastive Language-Image Pre-training (CLIP) excels in global alignment with language but exhibits limited sensitivity to spatial information, leading to strong performance in zero-shot classification tasks but underperformance in tasks requiring precise spatial understanding. Recent approaches have introduced Region-Language Alignment (RLA) to enhance CLIP's performance in dense multimodal tasks by aligning regional visual representations with corresponding text inputs. However, we find that CLIP ViTs fine-tuned with RLA suffer from notable loss in spatial awareness, which is crucial for dense prediction tasks. To address this, we propose the Spatial Correlation Distillation (SCD) framework, which preserves CLIP's inherent spatial structure and mitigates the above degradation. To further enhance spatial correlations, we introduce a lightweight Refiner that extracts refined correlations directly from CLIP before feeding them into SCD, based on an intriguing finding that CLIP naturally captures high-quality dense features. Together, these components form a robust distillation framework that enables CLIP ViTs to integrate both visual-language and visual-centric improvements, achieving state-of-the-art results across various open-vocabulary dense prediction benchmarks.
title Refining CLIP's Spatial Awareness: A Visual-Centric Perspective
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.02328