Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jiamian, Sun, Guohao, Wang, Pichao, Liu, Dongfang, Dianat, Sohail, Rabbani, Majid, Rao, Raghuveer, Tao, Zhiqiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910385809915904
author Wang, Jiamian
Sun, Guohao
Wang, Pichao
Liu, Dongfang
Dianat, Sohail
Rabbani, Majid
Rao, Raghuveer
Tao, Zhiqiang
author_facet Wang, Jiamian
Sun, Guohao
Wang, Pichao
Liu, Dongfang
Dianat, Sohail
Rabbani, Majid
Rao, Raghuveer
Tao, Zhiqiang
contents The increasing prevalence of video clips has sparked growing interest in text-video retrieval. Recent advances focus on establishing a joint embedding space for text and video, relying on consistent embedding representations to compute similarity. However, the text content in existing datasets is generally short and concise, making it hard to fully describe the redundant semantics of a video. Correspondingly, a single text embedding may be less expressive to capture the video embedding and empower the retrieval. In this study, we propose a new stochastic text modeling method T-MASS, i.e., text is modeled as a stochastic embedding, to enrich text embedding with a flexible and resilient semantic range, yielding a text mass. To be specific, we introduce a similarity-aware radius module to adapt the scale of the text mass upon the given text-video pairs. Plus, we design and develop a support text regularization to further control the text mass during the training. The inference pipeline is also tailored to fully exploit the text mass for accurate retrieval. Empirical evidence suggests that T-MASS not only effectively attracts relevant text-video pairs while distancing irrelevant ones, but also enables the determination of precise text embeddings for relevant pairs. Our experimental results show a substantial improvement of T-MASS over baseline (3% to 6.3% by R@1). Also, T-MASS achieves state-of-the-art performance on five benchmark datasets, including MSRVTT, LSMDC, DiDeMo, VATEX, and Charades.
format Preprint
id arxiv_https___arxiv_org_abs_2403_17998
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
Wang, Jiamian
Sun, Guohao
Wang, Pichao
Liu, Dongfang
Dianat, Sohail
Rabbani, Majid
Rao, Raghuveer
Tao, Zhiqiang
Computer Vision and Pattern Recognition
The increasing prevalence of video clips has sparked growing interest in text-video retrieval. Recent advances focus on establishing a joint embedding space for text and video, relying on consistent embedding representations to compute similarity. However, the text content in existing datasets is generally short and concise, making it hard to fully describe the redundant semantics of a video. Correspondingly, a single text embedding may be less expressive to capture the video embedding and empower the retrieval. In this study, we propose a new stochastic text modeling method T-MASS, i.e., text is modeled as a stochastic embedding, to enrich text embedding with a flexible and resilient semantic range, yielding a text mass. To be specific, we introduce a similarity-aware radius module to adapt the scale of the text mass upon the given text-video pairs. Plus, we design and develop a support text regularization to further control the text mass during the training. The inference pipeline is also tailored to fully exploit the text mass for accurate retrieval. Empirical evidence suggests that T-MASS not only effectively attracts relevant text-video pairs while distancing irrelevant ones, but also enables the determination of precise text embeddings for relevant pairs. Our experimental results show a substantial improvement of T-MASS over baseline (3% to 6.3% by R@1). Also, T-MASS achieves state-of-the-art performance on five benchmark datasets, including MSRVTT, LSMDC, DiDeMo, VATEX, and Charades.
title Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.17998