Saved in:
Bibliographic Details
Main Authors: Zheng, Guangcong, Yuan, Jianlong, Wang, Bo, Huang, Haoyang, Ma, Guoqing, Duan, Nan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.20827
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915307150376960
author Zheng, Guangcong
Yuan, Jianlong
Wang, Bo
Huang, Haoyang
Ma, Guoqing
Duan, Nan
author_facet Zheng, Guangcong
Yuan, Jianlong
Wang, Bo
Huang, Haoyang
Ma, Guoqing
Duan, Nan
contents Generating long videos that can show complex stories, like movie scenes from scripts, has great promise and offers much more than short clips. However, current methods that use autoregression with diffusion models often struggle because their step-by-step process naturally leads to a serious error accumulation (drift). Also, many existing ways to make long videos focus on single, continuous scenes, making them less useful for stories with many events and changes. This paper introduces a new approach to solve these problems. First, we propose a novel way to annotate datasets at the frame-level, providing detailed text guidance needed for making complex, multi-scene long videos. This detailed guidance works with a Frame-Level Attention Mechanism to make sure text and video match precisely. A key feature is that each part (frame) within these windows can be guided by its own distinct text prompt. Our training uses Diffusion Forcing to provide the model with the ability to handle time flexibly. We tested our approach on difficult VBench 2.0 benchmarks ("Complex Plots" and "Complex Landscapes") based on the WanX2.1-T2V-1.3B model. The results show our method is better at following instructions in complex, changing scenes and creates high-quality long videos. We plan to share our dataset annotation methods and trained models with the research community. Project page: https://zgctroy.github.io/frame-level-captions .
format Preprint
id arxiv_https___arxiv_org_abs_2505_20827
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Frame-Level Captions for Long Video Generation with Complex Multi Scenes
Zheng, Guangcong
Yuan, Jianlong
Wang, Bo
Huang, Haoyang
Ma, Guoqing
Duan, Nan
Computer Vision and Pattern Recognition
Generating long videos that can show complex stories, like movie scenes from scripts, has great promise and offers much more than short clips. However, current methods that use autoregression with diffusion models often struggle because their step-by-step process naturally leads to a serious error accumulation (drift). Also, many existing ways to make long videos focus on single, continuous scenes, making them less useful for stories with many events and changes. This paper introduces a new approach to solve these problems. First, we propose a novel way to annotate datasets at the frame-level, providing detailed text guidance needed for making complex, multi-scene long videos. This detailed guidance works with a Frame-Level Attention Mechanism to make sure text and video match precisely. A key feature is that each part (frame) within these windows can be guided by its own distinct text prompt. Our training uses Diffusion Forcing to provide the model with the ability to handle time flexibly. We tested our approach on difficult VBench 2.0 benchmarks ("Complex Plots" and "Complex Landscapes") based on the WanX2.1-T2V-1.3B model. The results show our method is better at following instructions in complex, changing scenes and creates high-quality long videos. We plan to share our dataset annotation methods and trained models with the research community. Project page: https://zgctroy.github.io/frame-level-captions .
title Frame-Level Captions for Long Video Generation with Complex Multi Scenes
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.20827