CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Xi, Dianbing, Wang, Jiepeng, Liang, Yuanzhi, Qiu, Xi, Liu, Jialun, Pan, Hao, Huo, Yuchi, Wang, Rui, Huang, Haibin, Zhang, Chi, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
Inverse Rendering using Multi-Bounce Path Tracing and Reservoir Sampling
by: Dai, Yuxin, et al.
Published: (2024)
by: Dai, Yuxin, et al.
Published: (2024)
PFAvatar: Pose-Fusion 3D Personalized Avatar Reconstruction from Real-World Outfit-of-the-Day Photos
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
Fuse3D: Generating 3D Assets Controlled by Multi-Image Fusion
by: Jin, Xuancheng, et al.
Published: (2025)
by: Jin, Xuancheng, et al.
Published: (2025)
VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation
by: Zhang, Chi, et al.
Published: (2024)
by: Zhang, Chi, et al.
Published: (2024)
Controllable Weather Synthesis and Removal with Video Diffusion Models
by: Lin, Chih-Hao, et al.
Published: (2025)
by: Lin, Chih-Hao, et al.
Published: (2025)
SURF: Signature-Retained Fast Video Generation
by: Ding, Kaixin, et al.
Published: (2025)
by: Ding, Kaixin, et al.
Published: (2025)
VideoMat: Extracting PBR Materials from Video Diffusion Models
by: J. Munkberg, et al.
Published: (2025)
by: J. Munkberg, et al.
Published: (2025)
A Gpu-based solution for large-scale skeletal animation simulation
by: Pan, Xi
Published: (2025)
by: Pan, Xi
Published: (2025)
X2Video: Adapting Diffusion Models for Multimodal Controllable Neural Video Rendering
by: Huang, Zhitong, et al.
Published: (2025)
by: Huang, Zhitong, et al.
Published: (2025)
Real‐Time Polygonal Lighting of Iridescence Effect using Precomputed Monomial‐Gaussians
by: Zhengze Liu, et al.
Published: (2024)
by: Zhengze Liu, et al.
Published: (2024)
TransVDM: Motion-Constrained Video Diffusion Model for Transparent Video Synthesis
by: Li, Menghao, et al.
Published: (2025)
by: Li, Menghao, et al.
Published: (2025)
Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
by: Gu, Zekai, et al.
Published: (2025)
by: Gu, Zekai, et al.
Published: (2025)
VideoMat: Extracting PBR Materials from Video Diffusion Models
by: Munkberg, Jacob, et al.
Published: (2025)
by: Munkberg, Jacob, et al.
Published: (2025)
Cinematographic Camera Diffusion Model
by: Jiang, Hongda, et al.
Published: (2024)
by: Jiang, Hongda, et al.
Published: (2024)
VideoFrom3D: 3D Scene Video Generation via Complementary Image and Video Diffusion Models
by: Kim, Geonung, et al.
Published: (2025)
by: Kim, Geonung, et al.
Published: (2025)
Sketch-based Fluid Video Generation Using Motion-Guided Diffusion Models in Still Landscape Images
by: Jin, Hao, et al.
Published: (2025)
by: Jin, Hao, et al.
Published: (2025)
ComboStoc: Combinatorial Stochasticity for Diffusion Generative Models
by: Xu, Rui, et al.
Published: (2024)
by: Xu, Rui, et al.
Published: (2024)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
DiffusionRenderer: Neural Inverse and Forward Rendering with Video Diffusion Models
by: Liang, Ruofan, et al.
Published: (2025)
by: Liang, Ruofan, et al.
Published: (2025)
TopoCtrl: Post-Optimization Topology Editing Toward Target Structural Characteristics
by: Chen, Hongrui, et al.
Published: (2026)
by: Chen, Hongrui, et al.
Published: (2026)
DISK: Differentiable Sparse Kernel Complex for Efficient Spatially-Variant Convolution
by: Wu, Zhizhen, et al.
Published: (2025)
by: Wu, Zhizhen, et al.
Published: (2025)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
by: Kuang, Zhengfei, et al.
Published: (2024)
by: Kuang, Zhengfei, et al.
Published: (2024)
LDM: Large Tensorial SDF Model for Textured Mesh Generation
by: Xie, Rengan, et al.
Published: (2024)
by: Xie, Rengan, et al.
Published: (2024)
DragVideo: Interactive Drag-style Video Editing
by: Deng, Yufan, et al.
Published: (2023)
by: Deng, Yufan, et al.
Published: (2023)
Sketch Video Synthesis
by: Yudian Zheng, et al.
Published: (2024)
by: Yudian Zheng, et al.
Published: (2024)
Cinematographic Camera Diffusion Model
by: Hongda Jiang, et al.
Published: (2024)
by: Hongda Jiang, et al.
Published: (2024)
Controllable Video Generation: A Survey
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
MotionAgent: Fine-grained Controllable Video Generation via Motion Field Agent
by: Liao, Xinyao, et al.
Published: (2025)
by: Liao, Xinyao, et al.
Published: (2025)
SPAR: Automated Static Placement of Virtual Content via Ergonomic‐Aware Multi‐Objective Optimization
by: Yiwen Wang, et al.
Published: (2026)
by: Yiwen Wang, et al.
Published: (2026)
Persistent Homology-Driven Optimization of Effective Relative Density Range for Triply Periodic Minimal Surface
by: Depeng, Gao, et al.
Published: (2024)
by: Depeng, Gao, et al.
Published: (2024)
LuxDiT: Lighting Estimation with Video Diffusion Transformer
by: Liang, Ruofan, et al.
Published: (2025)
by: Liang, Ruofan, et al.
Published: (2025)
T2Bs: Text-to-Character Blendshapes via Video Generation
by: Luo, Jiahao, et al.
Published: (2025)
by: Luo, Jiahao, et al.
Published: (2025)
Video4DGen: Enhancing Video and 4D Generation through Mutual Optimization
by: Wang, Yikai, et al.
Published: (2025)
by: Wang, Yikai, et al.
Published: (2025)
TelePhysics: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
by: Cheng, Shihao, et al.
Published: (2026)
by: Cheng, Shihao, et al.
Published: (2026)
CrossGen: Learning and Generating Cross Fields for Quad Meshing
by: Dong, Qiujie, et al.
Published: (2025)
by: Dong, Qiujie, et al.
Published: (2025)
FreqPrior: Improving Video Diffusion Models with Frequency Filtering Gaussian Noise
by: Yuan, Yunlong, et al.
Published: (2025)
by: Yuan, Yunlong, et al.
Published: (2025)
Synchronized Multi‐Frame Diffusion for Temporally Consistent Video Stylization
by: Minshan Xie, et al.
Published: (2025)
by: Minshan Xie, et al.
Published: (2025)
ReLumix: Extending Image Relighting to Video via Video Diffusion Models
by: Wang, Lezhong, et al.
Published: (2025)
by: Wang, Lezhong, et al.
Published: (2025)
Similar Items
-
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
by: Xi, Dianbing, et al.
Published: (2025) -
Inverse Rendering using Multi-Bounce Path Tracing and Reservoir Sampling
by: Dai, Yuxin, et al.
Published: (2024) -
PFAvatar: Pose-Fusion 3D Personalized Avatar Reconstruction from Real-World Outfit-of-the-Day Photos
by: Xi, Dianbing, et al.
Published: (2025) -
Fuse3D: Generating 3D Assets Controlled by Multi-Image Fusion
by: Jin, Xuancheng, et al.
Published: (2025) -
VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation
by: Zhang, Chi, et al.
Published: (2024)