CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xi, Dianbing, Wang, Jiepeng, Liang, Yuanzhi, Qiu, Xi, Liu, Jialun, Pan, Hao, Huo, Yuchi, Wang, Rui, Huang, Haibin, Zhang, Chi, Li, Xuelong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912729988595712
author Xi, Dianbing
Wang, Jiepeng
Liang, Yuanzhi
Qiu, Xi
Liu, Jialun
Pan, Hao
Huo, Yuchi
Wang, Rui
Huang, Haibin
Zhang, Chi
Li, Xuelong
author_facet Xi, Dianbing
Wang, Jiepeng
Liang, Yuanzhi
Qiu, Xi
Liu, Jialun
Pan, Hao
Huo, Yuchi
Wang, Rui
Huang, Haibin
Zhang, Chi
Li, Xuelong
contents We tackle the dual challenges of video understanding and controllable video generation within a unified diffusion framework. Our key insights are two-fold: geometry-only cues (e.g., depth, edges) are insufficient: they specify layout but under-constrain appearance, materials, and illumination, limiting physically meaningful edits such as relighting or material swaps and often causing temporal drift. Enriching the model with additional graphics-based modalities (intrinsics and semantics) provides complementary constraints that both disambiguate understanding and enable precise, predictable control during generation. However, building a single model that uses many heterogeneous cues introduces two core difficulties. Architecturally, the model must accept any subset of modalities, remain robust to missing inputs, and inject control signals without sacrificing temporal consistency. Data-wise, training demands large-scale, temporally aligned supervision that ties real videos to per-pixel multimodal annotations. We then propose CtrlVDiff, a unified diffusion model trained with a Hybrid Modality Control Strategy (HMCS) that routes and fuses features from depth, normals, segmentation, edges, and graphics-based intrinsics (albedo, roughness, metallic), and re-renders videos from any chosen subset with strong temporal coherence. To enable this, we build MMVideo, a hybrid real-and-synthetic dataset aligned across modalities and captions. Across understanding and generation benchmarks, CtrlVDiff delivers superior controllability and fidelity, enabling layer-wise edits (relighting, material adjustment, object insertion) and surpassing state-of-the-art baselines while remaining robust when some modalities are unavailable.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21129
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
Xi, Dianbing
Wang, Jiepeng
Liang, Yuanzhi
Qiu, Xi
Liu, Jialun
Pan, Hao
Huo, Yuchi
Wang, Rui
Huang, Haibin
Zhang, Chi
Li, Xuelong
Computer Vision and Pattern Recognition
Graphics
We tackle the dual challenges of video understanding and controllable video generation within a unified diffusion framework. Our key insights are two-fold: geometry-only cues (e.g., depth, edges) are insufficient: they specify layout but under-constrain appearance, materials, and illumination, limiting physically meaningful edits such as relighting or material swaps and often causing temporal drift. Enriching the model with additional graphics-based modalities (intrinsics and semantics) provides complementary constraints that both disambiguate understanding and enable precise, predictable control during generation. However, building a single model that uses many heterogeneous cues introduces two core difficulties. Architecturally, the model must accept any subset of modalities, remain robust to missing inputs, and inject control signals without sacrificing temporal consistency. Data-wise, training demands large-scale, temporally aligned supervision that ties real videos to per-pixel multimodal annotations. We then propose CtrlVDiff, a unified diffusion model trained with a Hybrid Modality Control Strategy (HMCS) that routes and fuses features from depth, normals, segmentation, edges, and graphics-based intrinsics (albedo, roughness, metallic), and re-renders videos from any chosen subset with strong temporal coherence. To enable this, we build MMVideo, a hybrid real-and-synthetic dataset aligned across modalities and captions. Across understanding and generation benchmarks, CtrlVDiff delivers superior controllability and fidelity, enabling layer-wise edits (relighting, material adjustment, object insertion) and surpassing state-of-the-art baselines while remaining robust when some modalities are unavailable.
title CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2511.21129