Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Haowei, Shi, Yaya, Xu, Haiyang, Yuan, Chunfeng, Ye, Qinghao, Li, Chenliang, Yan, Ming, Zhang, Ji, Huang, Fei, Li, Bing, Hu, Weiming
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909120341213184
author Liu, Haowei
Shi, Yaya
Xu, Haiyang
Yuan, Chunfeng
Ye, Qinghao
Li, Chenliang
Yan, Ming
Zhang, Ji
Huang, Fei
Li, Bing
Hu, Weiming
author_facet Liu, Haowei
Shi, Yaya
Xu, Haiyang
Yuan, Chunfeng
Ye, Qinghao
Li, Chenliang
Yan, Ming
Zhang, Ji
Huang, Fei
Li, Bing
Hu, Weiming
contents In video-text retrieval, most existing methods adopt the dual-encoder architecture for fast retrieval, which employs two individual encoders to extract global latent representations for videos and texts. However, they face challenges in capturing fine-grained semantic concepts. In this work, we propose the UNIFY framework, which learns lexicon representations to capture fine-grained semantics and combines the strengths of latent and lexicon representations for video-text retrieval. Specifically, we map videos and texts into a pre-defined lexicon space, where each dimension corresponds to a semantic concept. A two-stage semantics grounding approach is proposed to activate semantically relevant dimensions and suppress irrelevant dimensions. The learned lexicon representations can thus reflect fine-grained semantics of videos and texts. Furthermore, to leverage the complementarity between latent and lexicon representations, we propose a unified learning scheme to facilitate mutual learning via structure sharing and self-distillation. Experimental results show our UNIFY framework largely outperforms previous video-text retrieval methods, with 4.8% and 8.2% Recall@1 improvement on MSR-VTT and DiDeMo respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2402_16769
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval
Liu, Haowei
Shi, Yaya
Xu, Haiyang
Yuan, Chunfeng
Ye, Qinghao
Li, Chenliang
Yan, Ming
Zhang, Ji
Huang, Fei
Li, Bing
Hu, Weiming
Computer Vision and Pattern Recognition
In video-text retrieval, most existing methods adopt the dual-encoder architecture for fast retrieval, which employs two individual encoders to extract global latent representations for videos and texts. However, they face challenges in capturing fine-grained semantic concepts. In this work, we propose the UNIFY framework, which learns lexicon representations to capture fine-grained semantics and combines the strengths of latent and lexicon representations for video-text retrieval. Specifically, we map videos and texts into a pre-defined lexicon space, where each dimension corresponds to a semantic concept. A two-stage semantics grounding approach is proposed to activate semantically relevant dimensions and suppress irrelevant dimensions. The learned lexicon representations can thus reflect fine-grained semantics of videos and texts. Furthermore, to leverage the complementarity between latent and lexicon representations, we propose a unified learning scheme to facilitate mutual learning via structure sharing and self-distillation. Experimental results show our UNIFY framework largely outperforms previous video-text retrieval methods, with 4.8% and 8.2% Recall@1 improvement on MSR-VTT and DiDeMo respectively.
title Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.16769