Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909120341213184 |
|---|---|
| author | Liu, Haowei Shi, Yaya Xu, Haiyang Yuan, Chunfeng Ye, Qinghao Li, Chenliang Yan, Ming Zhang, Ji Huang, Fei Li, Bing Hu, Weiming |
| author_facet | Liu, Haowei Shi, Yaya Xu, Haiyang Yuan, Chunfeng Ye, Qinghao Li, Chenliang Yan, Ming Zhang, Ji Huang, Fei Li, Bing Hu, Weiming |
| contents | In video-text retrieval, most existing methods adopt the dual-encoder architecture for fast retrieval, which employs two individual encoders to extract global latent representations for videos and texts. However, they face challenges in capturing fine-grained semantic concepts. In this work, we propose the UNIFY framework, which learns lexicon representations to capture fine-grained semantics and combines the strengths of latent and lexicon representations for video-text retrieval. Specifically, we map videos and texts into a pre-defined lexicon space, where each dimension corresponds to a semantic concept. A two-stage semantics grounding approach is proposed to activate semantically relevant dimensions and suppress irrelevant dimensions. The learned lexicon representations can thus reflect fine-grained semantics of videos and texts. Furthermore, to leverage the complementarity between latent and lexicon representations, we propose a unified learning scheme to facilitate mutual learning via structure sharing and self-distillation. Experimental results show our UNIFY framework largely outperforms previous video-text retrieval methods, with 4.8% and 8.2% Recall@1 improvement on MSR-VTT and DiDeMo respectively. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_16769 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval Liu, Haowei Shi, Yaya Xu, Haiyang Yuan, Chunfeng Ye, Qinghao Li, Chenliang Yan, Ming Zhang, Ji Huang, Fei Li, Bing Hu, Weiming Computer Vision and Pattern Recognition In video-text retrieval, most existing methods adopt the dual-encoder architecture for fast retrieval, which employs two individual encoders to extract global latent representations for videos and texts. However, they face challenges in capturing fine-grained semantic concepts. In this work, we propose the UNIFY framework, which learns lexicon representations to capture fine-grained semantics and combines the strengths of latent and lexicon representations for video-text retrieval. Specifically, we map videos and texts into a pre-defined lexicon space, where each dimension corresponds to a semantic concept. A two-stage semantics grounding approach is proposed to activate semantically relevant dimensions and suppress irrelevant dimensions. The learned lexicon representations can thus reflect fine-grained semantics of videos and texts. Furthermore, to leverage the complementarity between latent and lexicon representations, we propose a unified learning scheme to facilitate mutual learning via structure sharing and self-distillation. Experimental results show our UNIFY framework largely outperforms previous video-text retrieval methods, with 4.8% and 8.2% Recall@1 improvement on MSR-VTT and DiDeMo respectively. |
| title | Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2402.16769 |