Consistency-Preserving Diverse Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Xinshuang, Li, Runfa Blark, Nguyen, Truong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908837400805376
author Liu, Xinshuang
Li, Runfa Blark
Nguyen, Truong
author_facet Liu, Xinshuang
Li, Runfa Blark
Nguyen, Truong
contents Text-to-video generation is expensive, so only a few samples are typically produced per prompt. In this low-sample regime, maximizing the value of each batch requires high cross-video diversity. Recent methods improve diversity for image generation, but for videos they often degrade within-video temporal consistency and require costly backpropagation through a video decoder. We propose a joint-sampling framework for flow-matching video generators that improves batch diversity while preserving temporal consistency. Our approach applies diversity-driven updates and then removes only the components that would decrease a temporal-consistency objective. To avoid image-space gradients, we compute both objectives with lightweight latent-space models, avoiding video decoding and decoder backpropagation. Experiments on a state-of-the-art text-to-video flow-matching model show diversity comparable to strong joint-sampling baselines while substantially improving temporal consistency and color naturalness. Code will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15287
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Consistency-Preserving Diverse Video Generation
Liu, Xinshuang
Li, Runfa Blark
Nguyen, Truong
Computer Vision and Pattern Recognition
Text-to-video generation is expensive, so only a few samples are typically produced per prompt. In this low-sample regime, maximizing the value of each batch requires high cross-video diversity. Recent methods improve diversity for image generation, but for videos they often degrade within-video temporal consistency and require costly backpropagation through a video decoder. We propose a joint-sampling framework for flow-matching video generators that improves batch diversity while preserving temporal consistency. Our approach applies diversity-driven updates and then removes only the components that would decrease a temporal-consistency objective. To avoid image-space gradients, we compute both objectives with lightweight latent-space models, avoiding video decoding and decoder backpropagation. Experiments on a state-of-the-art text-to-video flow-matching model show diversity comparable to strong joint-sampling baselines while substantially improving temporal consistency and color naturalness. Code will be released.
title Consistency-Preserving Diverse Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.15287