ScalingNoise: Scaling Inference-Time Search for Generating Infinite Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Haolin, Tang, Feilong, Hu, Ming, Yin, Qingyu, Li, Yulong, Liu, Yexin, Peng, Zelin, Gao, Peng, He, Junjun, Ge, Zongyuan, Razzak, Imran
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913857011712000
author Yang, Haolin
Tang, Feilong
Hu, Ming
Yin, Qingyu
Li, Yulong
Liu, Yexin
Peng, Zelin
Gao, Peng
He, Junjun
Ge, Zongyuan
Razzak, Imran
author_facet Yang, Haolin
Tang, Feilong
Hu, Ming
Yin, Qingyu
Li, Yulong
Liu, Yexin
Peng, Zelin
Gao, Peng
He, Junjun
Ge, Zongyuan
Razzak, Imran
contents Video diffusion models (VDMs) facilitate the generation of high-quality videos, with current research predominantly concentrated on scaling efforts during training through improvements in data quality, computational resources, and model complexity. However, inference-time scaling has received less attention, with most approaches restricting models to a single generation attempt. Recent studies have uncovered the existence of "golden noises" that can enhance video quality during generation. Building on this, we find that guiding the scaling inference-time search of VDMs to identify better noise candidates not only evaluates the quality of the frames generated in the current step but also preserves the high-level object features by referencing the anchor frame from previous multi-chunks, thereby delivering long-term value. Our analysis reveals that diffusion models inherently possess flexible adjustments of computation by varying denoising steps, and even a one-step denoising approach, when guided by a reward signal, yields significant long-term benefits. Based on the observation, we proposeScalingNoise, a plug-and-play inference-time search strategy that identifies golden initial noises for the diffusion sampling process to improve global content consistency and visual diversity. Specifically, we perform one-step denoising to convert initial noises into a clip and subsequently evaluate its long-term value, leveraging a reward model anchored by previously generated content. Moreover, to preserve diversity, we sample candidates from a tilted noise distribution that up-weights promising noises. In this way, ScalingNoise significantly reduces noise-induced errors, ensuring more coherent and spatiotemporally consistent video generation. Extensive experiments on benchmark datasets demonstrate that the proposed ScalingNoise effectively improves long video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16400
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ScalingNoise: Scaling Inference-Time Search for Generating Infinite Videos
Yang, Haolin
Tang, Feilong
Hu, Ming
Yin, Qingyu
Li, Yulong
Liu, Yexin
Peng, Zelin
Gao, Peng
He, Junjun
Ge, Zongyuan
Razzak, Imran
Machine Learning
Video diffusion models (VDMs) facilitate the generation of high-quality videos, with current research predominantly concentrated on scaling efforts during training through improvements in data quality, computational resources, and model complexity. However, inference-time scaling has received less attention, with most approaches restricting models to a single generation attempt. Recent studies have uncovered the existence of "golden noises" that can enhance video quality during generation. Building on this, we find that guiding the scaling inference-time search of VDMs to identify better noise candidates not only evaluates the quality of the frames generated in the current step but also preserves the high-level object features by referencing the anchor frame from previous multi-chunks, thereby delivering long-term value. Our analysis reveals that diffusion models inherently possess flexible adjustments of computation by varying denoising steps, and even a one-step denoising approach, when guided by a reward signal, yields significant long-term benefits. Based on the observation, we proposeScalingNoise, a plug-and-play inference-time search strategy that identifies golden initial noises for the diffusion sampling process to improve global content consistency and visual diversity. Specifically, we perform one-step denoising to convert initial noises into a clip and subsequently evaluate its long-term value, leveraging a reward model anchored by previously generated content. Moreover, to preserve diversity, we sample candidates from a tilted noise distribution that up-weights promising noises. In this way, ScalingNoise significantly reduces noise-induced errors, ensuring more coherent and spatiotemporally consistent video generation. Extensive experiments on benchmark datasets demonstrate that the proposed ScalingNoise effectively improves long video generation.
title ScalingNoise: Scaling Inference-Time Search for Generating Infinite Videos
topic Machine Learning
url https://arxiv.org/abs/2503.16400