FastInit: Fast Noise Initialization for Temporally Consistent Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bai, Chengyu, Li, Yuming, Zhao, Zhongyu, Chen, Jintao, Jia, Peidong, She, Qi, Lu, Ming, Zhang, Shanghang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908479061491712
author Bai, Chengyu
Li, Yuming
Zhao, Zhongyu
Chen, Jintao
Jia, Peidong
She, Qi
Lu, Ming
Zhang, Shanghang
author_facet Bai, Chengyu
Li, Yuming
Zhao, Zhongyu
Chen, Jintao
Jia, Peidong
She, Qi
Lu, Ming
Zhang, Shanghang
contents Video generation has made significant strides with the development of diffusion models; however, achieving high temporal consistency remains a challenging task. Recently, FreeInit identified a training-inference gap and introduced a method to iteratively refine the initial noise during inference. However, iterative refinement significantly increases the computational cost associated with video generation. In this paper, we introduce FastInit, a fast noise initialization method that eliminates the need for iterative refinement. FastInit learns a Video Noise Prediction Network (VNPNet) that takes random noise and a text prompt as input, generating refined noise in a single forward pass. Therefore, FastInit greatly enhances the efficiency of video generation while achieving high temporal consistency across frames. To train the VNPNet, we create a large-scale dataset consisting of pairs of text prompts, random noise, and refined noise. Extensive experiments with various text-to-video models show that our method consistently improves the quality and temporal consistency of the generated videos. FastInit not only provides a substantial improvement in video generation but also offers a practical solution that can be applied directly during inference. The code and dataset will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16119
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FastInit: Fast Noise Initialization for Temporally Consistent Video Generation
Bai, Chengyu
Li, Yuming
Zhao, Zhongyu
Chen, Jintao
Jia, Peidong
She, Qi
Lu, Ming
Zhang, Shanghang
Computer Vision and Pattern Recognition
Artificial Intelligence
Video generation has made significant strides with the development of diffusion models; however, achieving high temporal consistency remains a challenging task. Recently, FreeInit identified a training-inference gap and introduced a method to iteratively refine the initial noise during inference. However, iterative refinement significantly increases the computational cost associated with video generation. In this paper, we introduce FastInit, a fast noise initialization method that eliminates the need for iterative refinement. FastInit learns a Video Noise Prediction Network (VNPNet) that takes random noise and a text prompt as input, generating refined noise in a single forward pass. Therefore, FastInit greatly enhances the efficiency of video generation while achieving high temporal consistency across frames. To train the VNPNet, we create a large-scale dataset consisting of pairs of text prompts, random noise, and refined noise. Extensive experiments with various text-to-video models show that our method consistently improves the quality and temporal consistency of the generated videos. FastInit not only provides a substantial improvement in video generation but also offers a practical solution that can be applied directly during inference. The code and dataset will be released.
title FastInit: Fast Noise Initialization for Temporally Consistent Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.16119