PRVR: Partially Relevant Video Retrieval

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Xianke, Liu, Daizong, Yang, Xun, Li, Xirong, Dong, Jianfeng, Wang, Meng, Wang, Xun
Formato: Preprint
Publicado: 2022
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908582248710144
author Chen, Xianke
Liu, Daizong
Yang, Xun
Li, Xirong
Dong, Jianfeng
Wang, Meng
Wang, Xun
author_facet Chen, Xianke
Liu, Daizong
Yang, Xun
Li, Xirong
Dong, Jianfeng
Wang, Meng
Wang, Xun
contents In current text-to-video retrieval (T2VR), videos to be retrieved have been properly trimmed so that a correspondence between the videos and ad-hoc textual queries naturally exists. Note in practice that videos circulated on the Internet and social media platforms, while being relatively short, are typically rich in their content. Often, multiple scenes / actions / events are shown in a single video, leading to a more challenging T2VR setting wherein only part of the video content is relevant w.r.t. a given query. This paper presents a first study on this setting which we term Partially Relevant Video Retrieval (PRVR). Considering that a video typically consists of multiple moments, a video is regarded as partially relevant w.r.t. to a given query if it contains a query-related moment. We formulate the PRVR task as a multiple instance learning problem, and propose a Multi-Scale Similarity Learning (MS-SL++) network that jointly learns both clip-scale and frame-scale similarities to determine the partial relevance between video-query pairs. Extensive experiments on three diverse video-text datasets (TVshow Retrieval, ActivityNet-Captions and Charades-STA) demonstrate the viability of the proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2208_12510
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle PRVR: Partially Relevant Video Retrieval
Chen, Xianke
Liu, Daizong
Yang, Xun
Li, Xirong
Dong, Jianfeng
Wang, Meng
Wang, Xun
Computer Vision and Pattern Recognition
Multimedia
In current text-to-video retrieval (T2VR), videos to be retrieved have been properly trimmed so that a correspondence between the videos and ad-hoc textual queries naturally exists. Note in practice that videos circulated on the Internet and social media platforms, while being relatively short, are typically rich in their content. Often, multiple scenes / actions / events are shown in a single video, leading to a more challenging T2VR setting wherein only part of the video content is relevant w.r.t. a given query. This paper presents a first study on this setting which we term Partially Relevant Video Retrieval (PRVR). Considering that a video typically consists of multiple moments, a video is regarded as partially relevant w.r.t. to a given query if it contains a query-related moment. We formulate the PRVR task as a multiple instance learning problem, and propose a Multi-Scale Similarity Learning (MS-SL++) network that jointly learns both clip-scale and frame-scale similarities to determine the partial relevance between video-query pairs. Extensive experiments on three diverse video-text datasets (TVshow Retrieval, ActivityNet-Captions and Charades-STA) demonstrate the viability of the proposed method.
title PRVR: Partially Relevant Video Retrieval
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2208.12510