OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xi, Dianbing, Wang, Jiepeng, Liang, Yuanzhi, Qiu, Xi, Huo, Yuchi, Wang, Rui, Zhang, Chi, Li, Xuelong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911267253387264
author Xi, Dianbing
Wang, Jiepeng
Liang, Yuanzhi
Qiu, Xi
Huo, Yuchi
Wang, Rui
Zhang, Chi
Li, Xuelong
author_facet Xi, Dianbing
Wang, Jiepeng
Liang, Yuanzhi
Qiu, Xi
Huo, Yuchi
Wang, Rui
Zhang, Chi
Li, Xuelong
contents In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual modalities in the color space to learn a joint distribution, while employing an adaptive control strategy that dynamically adjusts the role of each visual modality during the diffusion process, either as a generation modality or a conditioning modality. Our framework supports three key capabilities: (1) Text-conditioned video generation, where all modalities are jointly synthesized from a textual prompt; (2) Video understanding, where structural modalities are predicted from rgb inputs in a coherent manner; and (3) X-conditioned video generation, where video synthesis is guided by finegrained inputs such as depth, canny and segmentation. Extensive experiments demonstrate that OmniVDiff achieves state-of-the-art performance in video generation tasks and competitive results in video understanding. Its flexibility and scalability make it well-suited for downstream applications such as video-to-video translation, modality adaptation for visual tasks, and scene reconstruction.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10825
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
Xi, Dianbing
Wang, Jiepeng
Liang, Yuanzhi
Qiu, Xi
Huo, Yuchi
Wang, Rui
Zhang, Chi
Li, Xuelong
Computer Vision and Pattern Recognition
In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual modalities in the color space to learn a joint distribution, while employing an adaptive control strategy that dynamically adjusts the role of each visual modality during the diffusion process, either as a generation modality or a conditioning modality. Our framework supports three key capabilities: (1) Text-conditioned video generation, where all modalities are jointly synthesized from a textual prompt; (2) Video understanding, where structural modalities are predicted from rgb inputs in a coherent manner; and (3) X-conditioned video generation, where video synthesis is guided by finegrained inputs such as depth, canny and segmentation. Extensive experiments demonstrate that OmniVDiff achieves state-of-the-art performance in video generation tasks and competitive results in video understanding. Its flexibility and scalability make it well-suited for downstream applications such as video-to-video translation, modality adaptation for visual tasks, and scene reconstruction.
title OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.10825