Saliency-Guided DETR for Moment Retrieval and Highlight Detection

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gordeev, Aleksandr, Dokholyan, Vladimir, Tolstykh, Irina, Kuprashevich, Maksim
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908806467813376
author Gordeev, Aleksandr
Dokholyan, Vladimir
Tolstykh, Irina
Kuprashevich, Maksim
author_facet Gordeev, Aleksandr
Dokholyan, Vladimir
Tolstykh, Irina
Kuprashevich, Maksim
contents Existing approaches for video moment retrieval and highlight detection are not able to align text and video features efficiently, resulting in unsatisfying performance and limited production usage. To address this, we propose a novel architecture that utilizes recent foundational video models designed for such alignment. Combined with the introduced Saliency-Guided Cross Attention mechanism and a hybrid DETR architecture, our approach significantly enhances performance in both moment retrieval and highlight detection tasks. For even better improvement, we developed InterVid-MR, a large-scale and high-quality dataset for pretraining. Using it, our architecture achieves state-of-the-art results on the QVHighlights, Charades-STA and TACoS benchmarks. The proposed approach provides an efficient and scalable solution for both zero-shot and fine-tuning scenarios in video-language tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_01615
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Saliency-Guided DETR for Moment Retrieval and Highlight Detection
Gordeev, Aleksandr
Dokholyan, Vladimir
Tolstykh, Irina
Kuprashevich, Maksim
Computer Vision and Pattern Recognition
Existing approaches for video moment retrieval and highlight detection are not able to align text and video features efficiently, resulting in unsatisfying performance and limited production usage. To address this, we propose a novel architecture that utilizes recent foundational video models designed for such alignment. Combined with the introduced Saliency-Guided Cross Attention mechanism and a hybrid DETR architecture, our approach significantly enhances performance in both moment retrieval and highlight detection tasks. For even better improvement, we developed InterVid-MR, a large-scale and high-quality dataset for pretraining. Using it, our architecture achieves state-of-the-art results on the QVHighlights, Charades-STA and TACoS benchmarks. The proposed approach provides an efficient and scalable solution for both zero-shot and fine-tuning scenarios in video-language tasks.
title Saliency-Guided DETR for Moment Retrieval and Highlight Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.01615