Exploring Iterative Refinement with Diffusion Models for Video Grounding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liang, Xiao, Shi, Tao, Liang, Yaoyuan, Tao, Te, Huang, Shao-Lun
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910283078828032
author Liang, Xiao
Shi, Tao
Liang, Yaoyuan
Tao, Te
Huang, Shao-Lun
author_facet Liang, Xiao
Shi, Tao
Liang, Yaoyuan
Tao, Te
Huang, Shao-Lun
contents Video grounding aims to localize the target moment in an untrimmed video corresponding to a given sentence query. Existing methods typically select the best prediction from a set of predefined proposals or directly regress the target span in a single-shot manner, resulting in the absence of a systematical prediction refinement process. In this paper, we propose DiffusionVG, a novel framework with diffusion models that formulates video grounding as a conditional generation task, where the target span is generated from Gaussian noise inputs and interatively refined in the reverse diffusion process. During training, DiffusionVG progressively adds noise to the target span with a fixed forward diffusion process and learns to recover the target span in the reverse diffusion process. In inference, DiffusionVG can generate the target span from Gaussian noise inputs by the learned reverse diffusion process conditioned on the video-sentence representations. Without bells and whistles, our DiffusionVG demonstrates superior performance compared to existing well-crafted models on mainstream Charades-STA, ActivityNet Captions and TACoS benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2310_17189
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Exploring Iterative Refinement with Diffusion Models for Video Grounding
Liang, Xiao
Shi, Tao
Liang, Yaoyuan
Tao, Te
Huang, Shao-Lun
Computer Vision and Pattern Recognition
Video grounding aims to localize the target moment in an untrimmed video corresponding to a given sentence query. Existing methods typically select the best prediction from a set of predefined proposals or directly regress the target span in a single-shot manner, resulting in the absence of a systematical prediction refinement process. In this paper, we propose DiffusionVG, a novel framework with diffusion models that formulates video grounding as a conditional generation task, where the target span is generated from Gaussian noise inputs and interatively refined in the reverse diffusion process. During training, DiffusionVG progressively adds noise to the target span with a fixed forward diffusion process and learns to recover the target span in the reverse diffusion process. In inference, DiffusionVG can generate the target span from Gaussian noise inputs by the learned reverse diffusion process conditioned on the video-sentence representations. Without bells and whistles, our DiffusionVG demonstrates superior performance compared to existing well-crafted models on mainstream Charades-STA, ActivityNet Captions and TACoS benchmarks.
title Exploring Iterative Refinement with Diffusion Models for Video Grounding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2310.17189