LongDiff: Training-Free Long Video Generation in One Go

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhuoling, Rahmani, Hossein, Ke, Qiuhong, Liu, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916660487651328
author Li, Zhuoling
Rahmani, Hossein
Ke, Qiuhong
Liu, Jun
author_facet Li, Zhuoling
Rahmani, Hossein
Ke, Qiuhong
Liu, Jun
contents Video diffusion models have recently achieved remarkable results in video generation. Despite their encouraging performance, most of these models are mainly designed and trained for short video generation, leading to challenges in maintaining temporal consistency and visual details in long video generation. In this paper, we propose LongDiff, a novel training-free method consisting of carefully designed components \ -- Position Mapping (PM) and Informative Frame Selection (IFS) \ -- to tackle two key challenges that hinder short-to-long video generation generalization: temporal position ambiguity and information dilution. Our LongDiff unlocks the potential of off-the-shelf video diffusion models to achieve high-quality long video generation in one go. Extensive experiments demonstrate the efficacy of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18150
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LongDiff: Training-Free Long Video Generation in One Go
Li, Zhuoling
Rahmani, Hossein
Ke, Qiuhong
Liu, Jun
Computer Vision and Pattern Recognition
Video diffusion models have recently achieved remarkable results in video generation. Despite their encouraging performance, most of these models are mainly designed and trained for short video generation, leading to challenges in maintaining temporal consistency and visual details in long video generation. In this paper, we propose LongDiff, a novel training-free method consisting of carefully designed components \ -- Position Mapping (PM) and Informative Frame Selection (IFS) \ -- to tackle two key challenges that hinder short-to-long video generation generalization: temporal position ambiguity and information dilution. Our LongDiff unlocks the potential of off-the-shelf video diffusion models to achieve high-quality long video generation in one go. Extensive experiments demonstrate the efficacy of our method.
title LongDiff: Training-Free Long Video Generation in One Go
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.18150