Multi-sentence Video Grounding for Long Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Wei, Wang, Xin, Chen, Hong, Zhang, Zeyang, Zhu, Wenwu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913435162247168
author Feng, Wei
Wang, Xin
Chen, Hong
Zhang, Zeyang
Zhu, Wenwu
author_facet Feng, Wei
Wang, Xin
Chen, Hong
Zhang, Zeyang
Zhu, Wenwu
contents Video generation has witnessed great success recently, but their application in generating long videos still remains challenging due to the difficulty in maintaining the temporal consistency of generated videos and the high memory cost during generation. To tackle the problems, in this paper, we propose a brave and new idea of Multi-sentence Video Grounding for Long Video Generation, connecting the massive video moment retrieval to the video generation task for the first time, providing a new paradigm for long video generation. The method of our work can be summarized as three steps: (i) We design sequential scene text prompts as the queries for video grounding, utilizing the massive video moment retrieval to search for video moment segments that meet the text requirements in the video database. (ii) Based on the source frames of retrieved video moment segments, we adopt video editing methods to create new video content while preserving the temporal consistency of the retrieved video. Since the editing can be conducted segment by segment, and even frame by frame, it largely reduces the memory cost. (iii) We also attempt video morphing and personalized generation methods to improve the subject consistency of long video generation, providing ablation experimental results for the subtasks of long video generation. Our approach seamlessly extends the development in image/video editing, video morphing and personalized generation, and video grounding to the long video generation, offering effective solutions for generating long videos at low memory cost.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13219
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-sentence Video Grounding for Long Video Generation
Feng, Wei
Wang, Xin
Chen, Hong
Zhang, Zeyang
Zhu, Wenwu
Computer Vision and Pattern Recognition
Video generation has witnessed great success recently, but their application in generating long videos still remains challenging due to the difficulty in maintaining the temporal consistency of generated videos and the high memory cost during generation. To tackle the problems, in this paper, we propose a brave and new idea of Multi-sentence Video Grounding for Long Video Generation, connecting the massive video moment retrieval to the video generation task for the first time, providing a new paradigm for long video generation. The method of our work can be summarized as three steps: (i) We design sequential scene text prompts as the queries for video grounding, utilizing the massive video moment retrieval to search for video moment segments that meet the text requirements in the video database. (ii) Based on the source frames of retrieved video moment segments, we adopt video editing methods to create new video content while preserving the temporal consistency of the retrieved video. Since the editing can be conducted segment by segment, and even frame by frame, it largely reduces the memory cost. (iii) We also attempt video morphing and personalized generation methods to improve the subject consistency of long video generation, providing ablation experimental results for the subtasks of long video generation. Our approach seamlessly extends the development in image/video editing, video morphing and personalized generation, and video grounding to the long video generation, offering effective solutions for generating long videos at low memory cost.
title Multi-sentence Video Grounding for Long Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.13219