Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Moon, WonJun, Cho, Cheol-Ho, Jun, Woojin, Shim, Minho, Kim, Taeoh, Lee, Inwoong, Wee, Dongyoon, Heo, Jae-Pil
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909583278080000
author Moon, WonJun
Cho, Cheol-Ho
Jun, Woojin
Shim, Minho
Kim, Taeoh
Lee, Inwoong
Wee, Dongyoon
Heo, Jae-Pil
author_facet Moon, WonJun
Cho, Cheol-Ho
Jun, Woojin
Shim, Minho
Kim, Taeoh
Lee, Inwoong
Wee, Dongyoon
Heo, Jae-Pil
contents In a retrieval system, simultaneously achieving search accuracy and efficiency is inherently challenging. This challenge is particularly pronounced in partially relevant video retrieval (PRVR), where incorporating more diverse context representations at varying temporal scales for each video enhances accuracy but increases computational and memory costs. To address this dichotomy, we propose a prototypical PRVR framework that encodes diverse contexts within a video into a fixed number of prototypes. We then introduce several strategies to enhance text association and video understanding within the prototypes, along with an orthogonal objective to ensure that the prototypes capture a diverse range of content. To keep the prototypes searchable via text queries while accurately encoding video contexts, we implement cross- and uni-modal reconstruction tasks. The cross-modal reconstruction task aligns the prototypes with textual features within a shared space, while the uni-modal reconstruction task preserves all video contexts during encoding. Additionally, we employ a video mixing technique to provide weak guidance to further align prototypes and associated textual representations. Extensive evaluations on TVR, ActivityNet-Captions, and QVHighlights validate the effectiveness of our approach without sacrificing efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13035
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval
Moon, WonJun
Cho, Cheol-Ho
Jun, Woojin
Shim, Minho
Kim, Taeoh
Lee, Inwoong
Wee, Dongyoon
Heo, Jae-Pil
Computer Vision and Pattern Recognition
Artificial Intelligence
In a retrieval system, simultaneously achieving search accuracy and efficiency is inherently challenging. This challenge is particularly pronounced in partially relevant video retrieval (PRVR), where incorporating more diverse context representations at varying temporal scales for each video enhances accuracy but increases computational and memory costs. To address this dichotomy, we propose a prototypical PRVR framework that encodes diverse contexts within a video into a fixed number of prototypes. We then introduce several strategies to enhance text association and video understanding within the prototypes, along with an orthogonal objective to ensure that the prototypes capture a diverse range of content. To keep the prototypes searchable via text queries while accurately encoding video contexts, we implement cross- and uni-modal reconstruction tasks. The cross-modal reconstruction task aligns the prototypes with textual features within a shared space, while the uni-modal reconstruction task preserves all video contexts during encoding. Additionally, we employ a video mixing technique to provide weak guidance to further align prototypes and associated textual representations. Extensive evaluations on TVR, ActivityNet-Captions, and QVHighlights validate the effectiveness of our approach without sacrificing efficiency.
title Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2504.13035