Beyond Caption-Based Queries for Video Moment Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pujol-Perich, David, Clapés, Albert, Damen, Dima, Escalera, Sergio, Wray, Michael
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910038869671936
author Pujol-Perich, David
Clapés, Albert
Damen, Dima
Escalera, Sergio
Wray, Michael
author_facet Pujol-Perich, David
Clapés, Albert
Damen, Dima
Escalera, Sergio
Wray, Michael
contents In this work, we investigate the degradation of existing VMR methods, particularly of DETR architectures, when trained on caption-based queries but evaluated on search queries. For this, we introduce three benchmarks by modifying the textual queries in three public VMR datasets -- i.e., HD-EPIC, YouCook2 and ActivityNet-Captions. Our analysis reveals two key generalization challenges: (i) A language gap, arising from the linguistic under-specification of search queries, and (ii) a multi-moment gap, caused by the shift from single-moment to multi-moment queries. We also identify a critical issue in these architectures -- an active decoder-query collapse -- as a primary cause of the poor generalization to multi-moment instances. We mitigate this issue with architectural modifications that effectively increase the number of active decoder queries. Extensive experiments demonstrate that our approach improves performance on search queries by up to 14.82% mAP_m, and up to 21.83% mAP_m on multi-moment search queries. The code, models and data are available in the project webpage: https://davidpujol.github.io/beyond-vmr/
format Preprint
id arxiv_https___arxiv_org_abs_2603_02363
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Caption-Based Queries for Video Moment Retrieval
Pujol-Perich, David
Clapés, Albert
Damen, Dima
Escalera, Sergio
Wray, Michael
Computer Vision and Pattern Recognition
In this work, we investigate the degradation of existing VMR methods, particularly of DETR architectures, when trained on caption-based queries but evaluated on search queries. For this, we introduce three benchmarks by modifying the textual queries in three public VMR datasets -- i.e., HD-EPIC, YouCook2 and ActivityNet-Captions. Our analysis reveals two key generalization challenges: (i) A language gap, arising from the linguistic under-specification of search queries, and (ii) a multi-moment gap, caused by the shift from single-moment to multi-moment queries. We also identify a critical issue in these architectures -- an active decoder-query collapse -- as a primary cause of the poor generalization to multi-moment instances. We mitigate this issue with architectural modifications that effectively increase the number of active decoder queries. Extensive experiments demonstrate that our approach improves performance on search queries by up to 14.82% mAP_m, and up to 21.83% mAP_m on multi-moment search queries. The code, models and data are available in the project webpage: https://davidpujol.github.io/beyond-vmr/
title Beyond Caption-Based Queries for Video Moment Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.02363