FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Jangho, Park, Geon Yeong, Kwon, Gihyun, Ye, Jong Chul
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914582493134848
author Park, Jangho
Park, Geon Yeong
Kwon, Gihyun
Ye, Jong Chul
author_facet Park, Jangho
Park, Geon Yeong
Kwon, Gihyun
Ye, Jong Chul
contents Extending the generation horizon of video diffusion models to long sequences remains a long-standing and important challenge. Existing training-free approaches fall into two categories: extensions of bidirectional models, which are tightly coupled to specific architectures and suffer from quality degradation over long horizons, and autoregressive models, which accumulate drift errors due to exposure bias and tend to produce repetitive motion patterns. To address these issues, we propose a novel but simple inference-time approach for long video generation that is architecture-agnostic and requires no additional training. Our method generates long videos via overlapping sliding windows, where predicted clean samples from adjacent windows are blended via \emph{Tweedie matching} to enforce both \textbf{manifold constraint and temporal consistency} across overlap regions. \emph{Stochastic early-phase sampling} then synchronizes per-window trajectories by injecting fresh noise after each Tweedie matching correction in the high-noise phase, before transitioning to deterministic ODE sampling to preserve fine-grained visual fidelity. Applied to various video generation models, our method generates videos several times longer than the native window length while outperforming both training-free and autoregressive baselines in temporal consistency and visual quality, and further extends to audio-video joint generation and text-to-3DGS without any fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2605_20910
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching
Park, Jangho
Park, Geon Yeong
Kwon, Gihyun
Ye, Jong Chul
Computer Vision and Pattern Recognition
Extending the generation horizon of video diffusion models to long sequences remains a long-standing and important challenge. Existing training-free approaches fall into two categories: extensions of bidirectional models, which are tightly coupled to specific architectures and suffer from quality degradation over long horizons, and autoregressive models, which accumulate drift errors due to exposure bias and tend to produce repetitive motion patterns. To address these issues, we propose a novel but simple inference-time approach for long video generation that is architecture-agnostic and requires no additional training. Our method generates long videos via overlapping sliding windows, where predicted clean samples from adjacent windows are blended via \emph{Tweedie matching} to enforce both \textbf{manifold constraint and temporal consistency} across overlap regions. \emph{Stochastic early-phase sampling} then synchronizes per-window trajectories by injecting fresh noise after each Tweedie matching correction in the high-noise phase, before transitioning to deterministic ODE sampling to preserve fine-grained visual fidelity. Applied to various video generation models, our method generates videos several times longer than the native window length while outperforming both training-free and autoregressive baselines in temporal consistency and visual quality, and further extends to audio-video joint generation and text-to-3DGS without any fine-tuning.
title FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.20910