Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Yunlong, Guo, Yuanfan, Wang, Chunwei, Xu, Hang, Zhang, Li
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909448626241536
author Yuan, Yunlong
Guo, Yuanfan
Wang, Chunwei
Xu, Hang
Zhang, Li
author_facet Yuan, Yunlong
Guo, Yuanfan
Wang, Chunwei
Xu, Hang
Zhang, Li
contents Recent advances in diffusion models have greatly improved text-driven video generation. However, training models for long video generation demands significant computational power and extensive data, leading most video diffusion models to be limited to a small number of frames. Existing training-free methods that attempt to generate long videos using pre-trained short video diffusion models often struggle with issues such as insufficient motion dynamics and degraded video fidelity. In this paper, we present Brick-Diffusion, a novel, training-free approach capable of generating long videos of arbitrary length. Our method introduces a brick-to-wall denoising strategy, where the latent is denoised in segments, with a stride applied in subsequent iterations. This process mimics the construction of a staggered brick wall, where each brick represents a denoised segment, enabling communication between frames and improving overall video quality. Through quantitative and qualitative evaluations, we demonstrate that Brick-Diffusion outperforms existing baseline methods in generating high-fidelity videos.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02741
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising
Yuan, Yunlong
Guo, Yuanfan
Wang, Chunwei
Xu, Hang
Zhang, Li
Computer Vision and Pattern Recognition
Recent advances in diffusion models have greatly improved text-driven video generation. However, training models for long video generation demands significant computational power and extensive data, leading most video diffusion models to be limited to a small number of frames. Existing training-free methods that attempt to generate long videos using pre-trained short video diffusion models often struggle with issues such as insufficient motion dynamics and degraded video fidelity. In this paper, we present Brick-Diffusion, a novel, training-free approach capable of generating long videos of arbitrary length. Our method introduces a brick-to-wall denoising strategy, where the latent is denoised in segments, with a stride applied in subsequent iterations. This process mimics the construction of a staggered brick wall, where each brick represents a denoised segment, enabling communication between frames and improving overall video quality. Through quantitative and qualitative evaluations, we demonstrate that Brick-Diffusion outperforms existing baseline methods in generating high-fidelity videos.
title Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.02741