DetailCLIP: Injecting Image Details into CLIP's Feature Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911613104160768 |
|---|---|
| author | Zhang, Zilun Shen, Cuifeng Shen, Yuan Zhou, Xinyu Xiong, Huixin Zhao, Tiancheng Yin, Jianwei |
| author_facet | Zhang, Zilun Shen, Cuifeng Shen, Yuan Zhou, Xinyu Xiong, Huixin Zhao, Tiancheng Yin, Jianwei |
| contents | Although CLIP-like Visual Language Models provide a functional joint feature space for image and text, due to the limitation of the CILP-like model's image input size (e.g., 224), subtle details are lost in the feature representation if we input high-resolution images (e.g., 2240). Our proposed framework addresses this issue by generating a single feature representation for a high-resolution image that retains image details from different scales while sharing the same semantic space as the original CLIP. An application scenario is remote sensing text-image retrieval, where targets (e.g., vehicles and ships) often appear at tiny scales. To achieve this, we develop a feature fusion model that relies on CLIP features extracted from a carefully designed image patch method, dubbed Complete Cover. This method ensures comprehensive coverage of objects across various scales and is weakly supervised by image-agnostic class prompted queries. We evaluate our framework's performance using real-world and synthetic datasets, demonstrating significant improvements in image retrieval tasks based on class prompted queries. To further showcase our framework's capability in detail retrieval, we introduce a CLEVR-like synthetic dataset, named CLVER-DS. This fully annotated dataset offers a controllable object scale, allowing for a more thorough examination of our approach's effectiveness.Our code is publicly available at https://github.com/zilunzhang/DetailCLIP |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2208_14649 |
| institution | arXiv |
| publishDate | 2022 |
| record_format | arxiv |
| spellingShingle | DetailCLIP: Injecting Image Details into CLIP's Feature Space Zhang, Zilun Shen, Cuifeng Shen, Yuan Zhou, Xinyu Xiong, Huixin Zhao, Tiancheng Yin, Jianwei Computer Vision and Pattern Recognition Although CLIP-like Visual Language Models provide a functional joint feature space for image and text, due to the limitation of the CILP-like model's image input size (e.g., 224), subtle details are lost in the feature representation if we input high-resolution images (e.g., 2240). Our proposed framework addresses this issue by generating a single feature representation for a high-resolution image that retains image details from different scales while sharing the same semantic space as the original CLIP. An application scenario is remote sensing text-image retrieval, where targets (e.g., vehicles and ships) often appear at tiny scales. To achieve this, we develop a feature fusion model that relies on CLIP features extracted from a carefully designed image patch method, dubbed Complete Cover. This method ensures comprehensive coverage of objects across various scales and is weakly supervised by image-agnostic class prompted queries. We evaluate our framework's performance using real-world and synthetic datasets, demonstrating significant improvements in image retrieval tasks based on class prompted queries. To further showcase our framework's capability in detail retrieval, we introduce a CLEVR-like synthetic dataset, named CLVER-DS. This fully annotated dataset offers a controllable object scale, allowing for a more thorough examination of our approach's effectiveness.Our code is publicly available at https://github.com/zilunzhang/DetailCLIP |
| title | DetailCLIP: Injecting Image Details into CLIP's Feature Space |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2208.14649 |