Beyond Simple Edits: Composed Video Retrieval with Dense Modifications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Thawakar, Omkar, Demidov, Dmitry, Thawkar, Ritesh, Anwer, Rao Muhammad, Shah, Mubarak, Khan, Fahad Shahbaz, Khan, Salman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909743714402304
author Thawakar, Omkar
Demidov, Dmitry
Thawkar, Ritesh
Anwer, Rao Muhammad
Shah, Mubarak
Khan, Fahad Shahbaz
Khan, Salman
author_facet Thawakar, Omkar
Demidov, Dmitry
Thawkar, Ritesh
Anwer, Rao Muhammad
Shah, Mubarak
Khan, Fahad Shahbaz
Khan, Salman
contents Composed video retrieval is a challenging task that strives to retrieve a target video based on a query video and a textual description detailing specific modifications. Standard retrieval frameworks typically struggle to handle the complexity of fine-grained compositional queries and variations in temporal understanding limiting their retrieval ability in the fine-grained setting. To address this issue, we introduce a novel dataset that captures both fine-grained and composed actions across diverse video segments, enabling more detailed compositional changes in retrieved video content. The proposed dataset, named Dense-WebVid-CoVR, consists of 1.6 million samples with dense modification text that is around seven times more than its existing counterpart. We further develop a new model that integrates visual and textual information through Cross-Attention (CA) fusion using grounded text encoder, enabling precise alignment between dense query modifications and target videos. The proposed model achieves state-of-the-art results surpassing existing methods on all metrics. Notably, it achieves 71.3\% Recall@1 in visual+text setting and outperforms the state-of-the-art by 3.4\%, highlighting its efficacy in terms of leveraging detailed video descriptions and dense modification texts. Our proposed dataset, code, and model are available at :https://github.com/OmkarThawakar/BSE-CoVR
format Preprint
id arxiv_https___arxiv_org_abs_2508_14039
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
Thawakar, Omkar
Demidov, Dmitry
Thawkar, Ritesh
Anwer, Rao Muhammad
Shah, Mubarak
Khan, Fahad Shahbaz
Khan, Salman
Computer Vision and Pattern Recognition
Composed video retrieval is a challenging task that strives to retrieve a target video based on a query video and a textual description detailing specific modifications. Standard retrieval frameworks typically struggle to handle the complexity of fine-grained compositional queries and variations in temporal understanding limiting their retrieval ability in the fine-grained setting. To address this issue, we introduce a novel dataset that captures both fine-grained and composed actions across diverse video segments, enabling more detailed compositional changes in retrieved video content. The proposed dataset, named Dense-WebVid-CoVR, consists of 1.6 million samples with dense modification text that is around seven times more than its existing counterpart. We further develop a new model that integrates visual and textual information through Cross-Attention (CA) fusion using grounded text encoder, enabling precise alignment between dense query modifications and target videos. The proposed model achieves state-of-the-art results surpassing existing methods on all metrics. Notably, it achieves 71.3\% Recall@1 in visual+text setting and outperforms the state-of-the-art by 3.4\%, highlighting its efficacy in terms of leveraging detailed video descriptions and dense modification texts. Our proposed dataset, code, and model are available at :https://github.com/OmkarThawakar/BSE-CoVR
title Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.14039