EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915303317831680 |
|---|---|
| author | Meng, GuangHao He, Sunan Wang, Jinpeng Dai, Tao Zhang, Letian Zhu, Jieming Li, Qing Wang, Gang Zhang, Rui Jiang, Yong |
| author_facet | Meng, GuangHao He, Sunan Wang, Jinpeng Dai, Tao Zhang, Letian Zhu, Jieming Li, Qing Wang, Gang Zhang, Rui Jiang, Yong |
| contents | Vision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries to retrieve corresponding images (or text). However, existing methods often neglect the rich visual semantics knowledge of entities, thus leading to incorrect retrieval results. To address this problem, we propose the Entity Visual Description enhanced CLIP (EvdCLIP), designed to leverage the visual knowledge of entities to enrich queries. Specifically, since humans recognize entities through visual cues, we employ a large language model (LLM) to generate Entity Visual Descriptions (EVDs) as alignment cues to complement textual data. These EVDs are then integrated into raw queries to create visually-rich, EVD-enhanced queries. Furthermore, recognizing that EVD-enhanced queries may introduce noise or low-quality expansions, we develop a novel, trainable EVD-aware Rewriter (EaRW) for vision-language retrieval tasks. EaRW utilizes EVD knowledge and the generative capabilities of the language model to effectively rewrite queries. With our specialized training strategy, EaRW can generate high-quality and low-noise EVD-enhanced queries. Extensive quantitative and qualitative experiments on image-text retrieval benchmarks validate the superiority of EvdCLIP on vision-language retrieval tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_18594 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models Meng, GuangHao He, Sunan Wang, Jinpeng Dai, Tao Zhang, Letian Zhu, Jieming Li, Qing Wang, Gang Zhang, Rui Jiang, Yong Computer Vision and Pattern Recognition Information Retrieval Vision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries to retrieve corresponding images (or text). However, existing methods often neglect the rich visual semantics knowledge of entities, thus leading to incorrect retrieval results. To address this problem, we propose the Entity Visual Description enhanced CLIP (EvdCLIP), designed to leverage the visual knowledge of entities to enrich queries. Specifically, since humans recognize entities through visual cues, we employ a large language model (LLM) to generate Entity Visual Descriptions (EVDs) as alignment cues to complement textual data. These EVDs are then integrated into raw queries to create visually-rich, EVD-enhanced queries. Furthermore, recognizing that EVD-enhanced queries may introduce noise or low-quality expansions, we develop a novel, trainable EVD-aware Rewriter (EaRW) for vision-language retrieval tasks. EaRW utilizes EVD knowledge and the generative capabilities of the language model to effectively rewrite queries. With our specialized training strategy, EaRW can generate high-quality and low-noise EVD-enhanced queries. Extensive quantitative and qualitative experiments on image-text retrieval benchmarks validate the superiority of EvdCLIP on vision-language retrieval tasks. |
| title | EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models |
| topic | Computer Vision and Pattern Recognition Information Retrieval |
| url | https://arxiv.org/abs/2505.18594 |