UniVid: Pyramid Diffusion Model for High Quality Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Xinyu, Yang, Binbin, Li, Tingtian, Yu, Yipeng, Lei, Sen
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912966024101888
author Xiao, Xinyu
Yang, Binbin
Li, Tingtian
Yu, Yipeng
Lei, Sen
author_facet Xiao, Xinyu
Yang, Binbin
Li, Tingtian
Yu, Yipeng
Lei, Sen
contents Diffusion-based text-to-video generation (T2V) or image-to-video (I2V) generation have emerged as a prominent research focus. However, there exists a challenge in integrating the two generative paradigms into a unified model. In this paper, we present a unified video generation model (UniVid) with hybrid conditions of the text prompt and reference image. Given these two available controls, our model can extract objects' appearance and their motion descriptions from textual prompts, while obtaining texture details and structural information from image clues to guide the video generation process. Specifically, we scale up the pre-trained text-to-image diffusion model for generating temporally coherent frames via introducing our temporal-pyramid cross-frame spatial-temporal attention modules and convolutions. To support bimodal control, we introduce a dual-stream cross-attention mechanism, whose attention scores can be freely re-weighted for interpolation of between single and two modalities controls during inference. Extensive experiments showcase that our UniVid achieves superior temporal coherence on T2V, I2V and (T+I)2V tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13739
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UniVid: Pyramid Diffusion Model for High Quality Video Generation
Xiao, Xinyu
Yang, Binbin
Li, Tingtian
Yu, Yipeng
Lei, Sen
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
Diffusion-based text-to-video generation (T2V) or image-to-video (I2V) generation have emerged as a prominent research focus. However, there exists a challenge in integrating the two generative paradigms into a unified model. In this paper, we present a unified video generation model (UniVid) with hybrid conditions of the text prompt and reference image. Given these two available controls, our model can extract objects' appearance and their motion descriptions from textual prompts, while obtaining texture details and structural information from image clues to guide the video generation process. Specifically, we scale up the pre-trained text-to-image diffusion model for generating temporally coherent frames via introducing our temporal-pyramid cross-frame spatial-temporal attention modules and convolutions. To support bimodal control, we introduce a dual-stream cross-attention mechanism, whose attention scores can be freely re-weighted for interpolation of between single and two modalities controls during inference. Extensive experiments showcase that our UniVid achieves superior temporal coherence on T2V, I2V and (T+I)2V tasks.
title UniVid: Pyramid Diffusion Model for High Quality Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2603.13739