OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911267253387264 |
|---|---|
| author | Xi, Dianbing Wang, Jiepeng Liang, Yuanzhi Qiu, Xi Huo, Yuchi Wang, Rui Zhang, Chi Li, Xuelong |
| author_facet | Xi, Dianbing Wang, Jiepeng Liang, Yuanzhi Qiu, Xi Huo, Yuchi Wang, Rui Zhang, Chi Li, Xuelong |
| contents | In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual modalities in the color space to learn a joint distribution, while employing an adaptive control strategy that dynamically adjusts the role of each visual modality during the diffusion process, either as a generation modality or a conditioning modality. Our framework supports three key capabilities: (1) Text-conditioned video generation, where all modalities are jointly synthesized from a textual prompt; (2) Video understanding, where structural modalities are predicted from rgb inputs in a coherent manner; and (3) X-conditioned video generation, where video synthesis is guided by finegrained inputs such as depth, canny and segmentation. Extensive experiments demonstrate that OmniVDiff achieves state-of-the-art performance in video generation tasks and competitive results in video understanding. Its flexibility and scalability make it well-suited for downstream applications such as video-to-video translation, modality adaptation for visual tasks, and scene reconstruction. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_10825 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding Xi, Dianbing Wang, Jiepeng Liang, Yuanzhi Qiu, Xi Huo, Yuchi Wang, Rui Zhang, Chi Li, Xuelong Computer Vision and Pattern Recognition In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual modalities in the color space to learn a joint distribution, while employing an adaptive control strategy that dynamically adjusts the role of each visual modality during the diffusion process, either as a generation modality or a conditioning modality. Our framework supports three key capabilities: (1) Text-conditioned video generation, where all modalities are jointly synthesized from a textual prompt; (2) Video understanding, where structural modalities are predicted from rgb inputs in a coherent manner; and (3) X-conditioned video generation, where video synthesis is guided by finegrained inputs such as depth, canny and segmentation. Extensive experiments demonstrate that OmniVDiff achieves state-of-the-art performance in video generation tasks and competitive results in video understanding. Its flexibility and scalability make it well-suited for downstream applications such as video-to-video translation, modality adaptation for visual tasks, and scene reconstruction. |
| title | OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2504.10825 |