Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909448626241536 |
|---|---|
| author | Yuan, Yunlong Guo, Yuanfan Wang, Chunwei Xu, Hang Zhang, Li |
| author_facet | Yuan, Yunlong Guo, Yuanfan Wang, Chunwei Xu, Hang Zhang, Li |
| contents | Recent advances in diffusion models have greatly improved text-driven video generation. However, training models for long video generation demands significant computational power and extensive data, leading most video diffusion models to be limited to a small number of frames. Existing training-free methods that attempt to generate long videos using pre-trained short video diffusion models often struggle with issues such as insufficient motion dynamics and degraded video fidelity. In this paper, we present Brick-Diffusion, a novel, training-free approach capable of generating long videos of arbitrary length. Our method introduces a brick-to-wall denoising strategy, where the latent is denoised in segments, with a stride applied in subsequent iterations. This process mimics the construction of a staggered brick wall, where each brick represents a denoised segment, enabling communication between frames and improving overall video quality. Through quantitative and qualitative evaluations, we demonstrate that Brick-Diffusion outperforms existing baseline methods in generating high-fidelity videos. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_02741 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising Yuan, Yunlong Guo, Yuanfan Wang, Chunwei Xu, Hang Zhang, Li Computer Vision and Pattern Recognition Recent advances in diffusion models have greatly improved text-driven video generation. However, training models for long video generation demands significant computational power and extensive data, leading most video diffusion models to be limited to a small number of frames. Existing training-free methods that attempt to generate long videos using pre-trained short video diffusion models often struggle with issues such as insufficient motion dynamics and degraded video fidelity. In this paper, we present Brick-Diffusion, a novel, training-free approach capable of generating long videos of arbitrary length. Our method introduces a brick-to-wall denoising strategy, where the latent is denoised in segments, with a stride applied in subsequent iterations. This process mimics the construction of a staggered brick wall, where each brick represents a denoised segment, enabling communication between frames and improving overall video quality. Through quantitative and qualitative evaluations, we demonstrate that Brick-Diffusion outperforms existing baseline methods in generating high-fidelity videos. |
| title | Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2501.02741 |