EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Meng, GuangHao, He, Sunan, Wang, Jinpeng, Dai, Tao, Zhang, Letian, Zhu, Jieming, Li, Qing, Wang, Gang, Zhang, Rui, Jiang, Yong
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915303317831680
author Meng, GuangHao
He, Sunan
Wang, Jinpeng
Dai, Tao
Zhang, Letian
Zhu, Jieming
Li, Qing
Wang, Gang
Zhang, Rui
Jiang, Yong
author_facet Meng, GuangHao
He, Sunan
Wang, Jinpeng
Dai, Tao
Zhang, Letian
Zhu, Jieming
Li, Qing
Wang, Gang
Zhang, Rui
Jiang, Yong
contents Vision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries to retrieve corresponding images (or text). However, existing methods often neglect the rich visual semantics knowledge of entities, thus leading to incorrect retrieval results. To address this problem, we propose the Entity Visual Description enhanced CLIP (EvdCLIP), designed to leverage the visual knowledge of entities to enrich queries. Specifically, since humans recognize entities through visual cues, we employ a large language model (LLM) to generate Entity Visual Descriptions (EVDs) as alignment cues to complement textual data. These EVDs are then integrated into raw queries to create visually-rich, EVD-enhanced queries. Furthermore, recognizing that EVD-enhanced queries may introduce noise or low-quality expansions, we develop a novel, trainable EVD-aware Rewriter (EaRW) for vision-language retrieval tasks. EaRW utilizes EVD knowledge and the generative capabilities of the language model to effectively rewrite queries. With our specialized training strategy, EaRW can generate high-quality and low-noise EVD-enhanced queries. Extensive quantitative and qualitative experiments on image-text retrieval benchmarks validate the superiority of EvdCLIP on vision-language retrieval tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18594
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models
Meng, GuangHao
He, Sunan
Wang, Jinpeng
Dai, Tao
Zhang, Letian
Zhu, Jieming
Li, Qing
Wang, Gang
Zhang, Rui
Jiang, Yong
Computer Vision and Pattern Recognition
Information Retrieval
Vision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries to retrieve corresponding images (or text). However, existing methods often neglect the rich visual semantics knowledge of entities, thus leading to incorrect retrieval results. To address this problem, we propose the Entity Visual Description enhanced CLIP (EvdCLIP), designed to leverage the visual knowledge of entities to enrich queries. Specifically, since humans recognize entities through visual cues, we employ a large language model (LLM) to generate Entity Visual Descriptions (EVDs) as alignment cues to complement textual data. These EVDs are then integrated into raw queries to create visually-rich, EVD-enhanced queries. Furthermore, recognizing that EVD-enhanced queries may introduce noise or low-quality expansions, we develop a novel, trainable EVD-aware Rewriter (EaRW) for vision-language retrieval tasks. EaRW utilizes EVD knowledge and the generative capabilities of the language model to effectively rewrite queries. With our specialized training strategy, EaRW can generate high-quality and low-noise EVD-enhanced queries. Extensive quantitative and qualitative experiments on image-text retrieval benchmarks validate the superiority of EvdCLIP on vision-language retrieval tasks.
title EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models
topic Computer Vision and Pattern Recognition
Information Retrieval
url https://arxiv.org/abs/2505.18594