LongDiff: Training-Free Long Video Generation in One Go
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916660487651328 |
|---|---|
| author | Li, Zhuoling Rahmani, Hossein Ke, Qiuhong Liu, Jun |
| author_facet | Li, Zhuoling Rahmani, Hossein Ke, Qiuhong Liu, Jun |
| contents | Video diffusion models have recently achieved remarkable results in video generation. Despite their encouraging performance, most of these models are mainly designed and trained for short video generation, leading to challenges in maintaining temporal consistency and visual details in long video generation. In this paper, we propose LongDiff, a novel training-free method consisting of carefully designed components \ -- Position Mapping (PM) and Informative Frame Selection (IFS) \ -- to tackle two key challenges that hinder short-to-long video generation generalization: temporal position ambiguity and information dilution. Our LongDiff unlocks the potential of off-the-shelf video diffusion models to achieve high-quality long video generation in one go. Extensive experiments demonstrate the efficacy of our method. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_18150 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LongDiff: Training-Free Long Video Generation in One Go Li, Zhuoling Rahmani, Hossein Ke, Qiuhong Liu, Jun Computer Vision and Pattern Recognition Video diffusion models have recently achieved remarkable results in video generation. Despite their encouraging performance, most of these models are mainly designed and trained for short video generation, leading to challenges in maintaining temporal consistency and visual details in long video generation. In this paper, we propose LongDiff, a novel training-free method consisting of carefully designed components \ -- Position Mapping (PM) and Informative Frame Selection (IFS) \ -- to tackle two key challenges that hinder short-to-long video generation generalization: temporal position ambiguity and information dilution. Our LongDiff unlocks the potential of off-the-shelf video diffusion models to achieve high-quality long video generation in one go. Extensive experiments demonstrate the efficacy of our method. |
| title | LongDiff: Training-Free Long Video Generation in One Go |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.18150 |