Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bao, Fan, Xiang, Chendong, Yue, Gang, He, Guande, Zhu, Hongzhou, Zheng, Kaiwen, Zhao, Min, Liu, Shilong, Wang, Yaole, Zhu, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Consistency Diffusion Bridge Models
by: He, Guande, et al.
Published: (2024)
by: He, Guande, et al.
Published: (2024)
Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model
by: Zhao, Min, et al.
Published: (2024)
by: Zhao, Min, et al.
Published: (2024)
Elucidating the Preconditioning in Consistency Distillation
by: Zheng, Kaiwen, et al.
Published: (2025)
by: Zheng, Kaiwen, et al.
Published: (2025)
Diffusion Bridge Implicit Models
by: Zheng, Kaiwen, et al.
Published: (2024)
by: Zheng, Kaiwen, et al.
Published: (2024)
RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers
by: Zhao, Min, et al.
Published: (2025)
by: Zhao, Min, et al.
Published: (2025)
Schrodinger Bridges Beat Diffusion Models on Text-to-Speech Synthesis
by: Chen, Zehua, et al.
Published: (2023)
by: Chen, Zehua, et al.
Published: (2023)
minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models
by: Zhao, Min, et al.
Published: (2026)
by: Zhao, Min, et al.
Published: (2026)
UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
by: Zhao, Min, et al.
Published: (2025)
by: Zhao, Min, et al.
Published: (2025)
Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian Surfels
by: Wang, Yikai, et al.
Published: (2024)
by: Wang, Yikai, et al.
Published: (2024)
Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
by: Zhao, Min, et al.
Published: (2026)
by: Zhao, Min, et al.
Published: (2026)
Vidarc: Embodied Video Diffusion Model for Closed-loop Control
by: Feng, Yao, et al.
Published: (2025)
by: Feng, Yao, et al.
Published: (2025)
Direct Discriminative Optimization: Your Likelihood-Based Visual Generative Model is Secretly a GAN Discriminator
by: Zheng, Kaiwen, et al.
Published: (2025)
by: Zheng, Kaiwen, et al.
Published: (2025)
UltraImage: Rethinking Resolution Extrapolation in Image Diffusion Transformers
by: Zhao, Min, et al.
Published: (2025)
by: Zhao, Min, et al.
Published: (2025)
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
by: Feng, Yao, et al.
Published: (2025)
by: Feng, Yao, et al.
Published: (2025)
Geometry-Aware Rotary Position Embedding for Consistent Video World Model
by: Xiang, Chendong, et al.
Published: (2026)
by: Xiang, Chendong, et al.
Published: (2026)
Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
by: Huang, Xun, et al.
Published: (2025)
by: Huang, Xun, et al.
Published: (2025)
Causality in Video Diffusers is Separable from Denoising
by: Bai, Xingjian, et al.
Published: (2026)
by: Bai, Xingjian, et al.
Published: (2026)
Improved Techniques for Maximum Likelihood Estimation for Diffusion ODEs
by: Zheng, Kaiwen, et al.
Published: (2023)
by: Zheng, Kaiwen, et al.
Published: (2023)
Aligning Diffusion Behaviors with Q-functions for Efficient Continuous Control
by: Chen, Huayu, et al.
Published: (2024)
by: Chen, Huayu, et al.
Published: (2024)
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
by: Wang, Wenjing, et al.
Published: (2023)
by: Wang, Wenjing, et al.
Published: (2023)
TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
High‐Temperature Tribological Performance of Short Bamboo Fiber‐Reinforced Polymer Composite
by: Geng Hou, et al.
Published: (2025)
by: Geng Hou, et al.
Published: (2025)
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
by: Zhu, Zixin, et al.
Published: (2024)
by: Zhu, Zixin, et al.
Published: (2024)
VoiceBridge: General Speech Restoration with One-step Latent Bridge Models
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Exploring Aleatoric Uncertainty in Object Detection via Vision Foundation Models
by: Cui, Peng, et al.
Published: (2024)
by: Cui, Peng, et al.
Published: (2024)
Noise Contrastive Alignment of Language Models with Explicit Rewards
by: Chen, Huayu, et al.
Published: (2024)
by: Chen, Huayu, et al.
Published: (2024)
Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency
by: Zheng, Kaiwen, et al.
Published: (2025)
by: Zheng, Kaiwen, et al.
Published: (2025)
ContextAnyone: Context-Aware Diffusion for Character-Consistent Text-to-Video Generation
by: Mai, Ziyang, et al.
Published: (2025)
by: Mai, Ziyang, et al.
Published: (2025)
A Survey: Spatiotemporal Consistency in Video Generation
by: Yin, Zhiyu, et al.
Published: (2025)
by: Yin, Zhiyu, et al.
Published: (2025)
Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation
by: Liu, Lingyu, et al.
Published: (2026)
by: Liu, Lingyu, et al.
Published: (2026)
AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation
by: Zhu, Jiayin, et al.
Published: (2025)
by: Zhu, Jiayin, et al.
Published: (2025)
RoboTransfer: Controllable Geometry-Consistent Video Diffusion for Manipulation Policy Transfer
by: Liu, Liu, et al.
Published: (2025)
by: Liu, Liu, et al.
Published: (2025)
ROIC-DM: Robust Text Inference and Classification via Diffusion Model
by: Yuan, Shilong, et al.
Published: (2024)
by: Yuan, Shilong, et al.
Published: (2024)
DCDM: Divide-and-Conquer Diffusion Models for Consistency-Preserving Video Generation
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
Ultrasound Image-to-Video Synthesis via Latent Dynamic Diffusion Models
by: Chen, Tingxiu, et al.
Published: (2025)
by: Chen, Tingxiu, et al.
Published: (2025)
T2VSafetyBench: Evaluating the Safety of Text-to-Video Generative Models
by: Miao, Yibo, et al.
Published: (2024)
by: Miao, Yibo, et al.
Published: (2024)
AI Powered High Quality Text to Video Generation with Enhanced Temporal Consistency
by: Patel, Piyushkumar
Published: (2025)
by: Patel, Piyushkumar
Published: (2025)
Consistent Human Image and Video Generation with Spatially Conditioned Diffusion
by: Cao, Mingdeng, et al.
Published: (2024)
by: Cao, Mingdeng, et al.
Published: (2024)
Similar Items
-
Consistency Diffusion Bridge Models
by: He, Guande, et al.
Published: (2024) -
Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model
by: Zhao, Min, et al.
Published: (2024) -
Elucidating the Preconditioning in Consistency Distillation
by: Zheng, Kaiwen, et al.
Published: (2025) -
Diffusion Bridge Implicit Models
by: Zheng, Kaiwen, et al.
Published: (2024) -
RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers
by: Zhao, Min, et al.
Published: (2025)