UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shi, Bowen, Zhao, Peisen, Wang, Zichen, Zhang, Yuhang, Wang, Yaoming, Li, Jin, Dai, Wenrui, Zou, Junni, Xiong, Hongkai, Tian, Qi, Zhang, Xiaopeng
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913566584471552
author Shi, Bowen
Zhao, Peisen
Wang, Zichen
Zhang, Yuhang
Wang, Yaoming
Li, Jin
Dai, Wenrui
Zou, Junni
Xiong, Hongkai
Tian, Qi
Zhang, Xiaopeng
author_facet Shi, Bowen
Zhao, Peisen
Wang, Zichen
Zhang, Yuhang
Wang, Yaoming
Li, Jin
Dai, Wenrui
Zou, Junni
Xiong, Hongkai
Tian, Qi
Zhang, Xiaopeng
contents Vision-language foundation models, represented by Contrastive Language-Image Pre-training (CLIP), have gained increasing attention for jointly understanding both vision and textual tasks. However, existing approaches primarily focus on training models to match global image representations with textual descriptions, thereby overlooking the critical alignment between local regions and corresponding text tokens. This paper extends CLIP with multi-granularity alignment. Notably, we deliberately construct a new dataset comprising pseudo annotations at various levels of granularities, encompassing image-level, region-level as well as pixel-level captions and tags. Accordingly, we develop a Unified Multi-Granularity learning framework, termed UMG-CLIP, which simultaneously empowers the model with versatile perception abilities across different levels of detail. With parameter efficient tuning, UMG-CLIP surpasses current widely used CLIP variants and achieves state-of-the-art performance on diverse image understanding benchmarks, including open-world recognition, retrieval, semantic segmentation, and panoptic segmentation tasks. We believe that UMG-CLIP represents a valuable advancement in vision-language foundation models. The code is available at https://github.com/lygsbw/UMG-CLIP.
format Preprint
id arxiv_https___arxiv_org_abs_2401_06397
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding
Shi, Bowen
Zhao, Peisen
Wang, Zichen
Zhang, Yuhang
Wang, Yaoming
Li, Jin
Dai, Wenrui
Zou, Junni
Xiong, Hongkai
Tian, Qi
Zhang, Xiaopeng
Computer Vision and Pattern Recognition
Vision-language foundation models, represented by Contrastive Language-Image Pre-training (CLIP), have gained increasing attention for jointly understanding both vision and textual tasks. However, existing approaches primarily focus on training models to match global image representations with textual descriptions, thereby overlooking the critical alignment between local regions and corresponding text tokens. This paper extends CLIP with multi-granularity alignment. Notably, we deliberately construct a new dataset comprising pseudo annotations at various levels of granularities, encompassing image-level, region-level as well as pixel-level captions and tags. Accordingly, we develop a Unified Multi-Granularity learning framework, termed UMG-CLIP, which simultaneously empowers the model with versatile perception abilities across different levels of detail. With parameter efficient tuning, UMG-CLIP surpasses current widely used CLIP variants and achieves state-of-the-art performance on diverse image understanding benchmarks, including open-world recognition, retrieval, semantic segmentation, and panoptic segmentation tasks. We believe that UMG-CLIP represents a valuable advancement in vision-language foundation models. The code is available at https://github.com/lygsbw/UMG-CLIP.
title UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.06397