MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Guile, Huang, David, Bai, Dongfeng, Liu, Bingbing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ArmGS: Composite Gaussian Appearance Refinement for Modeling Dynamic Urban Environments
by: Wu, Guile, et al.
Published: (2025)
by: Wu, Guile, et al.
Published: (2025)
Language and Geometry Grounded Sparse Voxel Representations for Holistic Scene Understanding
by: Wu, Guile, et al.
Published: (2026)
by: Wu, Guile, et al.
Published: (2026)
Nighttime Autonomous Driving Scene Reconstruction with Physically-Based Gaussian Splatting
by: Kim, Tae-Kyeong, et al.
Published: (2026)
by: Kim, Tae-Kyeong, et al.
Published: (2026)
TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention
by: Huang, David, et al.
Published: (2026)
by: Huang, David, et al.
Published: (2026)
Efficient Depth-Guided Urban View Synthesis
by: Miao, Sheng, et al.
Published: (2024)
by: Miao, Sheng, et al.
Published: (2024)
EVolSplat: Efficient Volume-based Gaussian Splatting for Urban View Synthesis
by: Miao, Sheng, et al.
Published: (2025)
by: Miao, Sheng, et al.
Published: (2025)
DiVE: Efficient Multi-View Driving Scenes Generation Based on Video Diffusion Transformer
by: Jiang, Junpeng, et al.
Published: (2025)
by: Jiang, Junpeng, et al.
Published: (2025)
EVolSplat4D: Efficient Volume-based Gaussian Splatting for 4D Urban Scene Synthesis
by: Miao, Sheng, et al.
Published: (2026)
by: Miao, Sheng, et al.
Published: (2026)
Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention
by: Lu, Hannan, et al.
Published: (2024)
by: Lu, Hannan, et al.
Published: (2024)
VQA-Diff: Exploiting VQA and Diffusion for Zero-Shot Image-to-3D Vehicle Asset Generation in Autonomous Driving
by: Liu, Yibo, et al.
Published: (2024)
by: Liu, Yibo, et al.
Published: (2024)
UniScale: Unified Scale-Aware 3D Reconstruction for Multi-View Understanding via Prior Injection for Robotic Perception
by: Mahdavian, Mohammad, et al.
Published: (2026)
by: Mahdavian, Mohammad, et al.
Published: (2026)
AutoSplat: Constrained Gaussian Splatting for Autonomous Driving Scene Reconstruction
by: Khan, Mustafa, et al.
Published: (2024)
by: Khan, Mustafa, et al.
Published: (2024)
UniGaussian: Driving Scene Reconstruction from Multiple Camera Models via Unified Gaussian Representations
by: Ren, Yuan, et al.
Published: (2024)
by: Ren, Yuan, et al.
Published: (2024)
HUGS: Holistic Urban 3D Scene Understanding via Gaussian Splatting
by: Zhou, Hongyu, et al.
Published: (2024)
by: Zhou, Hongyu, et al.
Published: (2024)
MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second
by: Lin, Chenguo, et al.
Published: (2025)
by: Lin, Chenguo, et al.
Published: (2025)
SceneCrafter: Controllable Multi-View Driving Scene Editing
by: Zhu, Zehao, et al.
Published: (2025)
by: Zhu, Zehao, et al.
Published: (2025)
DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes
by: Mei, Jianbiao, et al.
Published: (2024)
by: Mei, Jianbiao, et al.
Published: (2024)
MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding
by: Bai, Purui, et al.
Published: (2026)
by: Bai, Purui, et al.
Published: (2026)
LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation
by: Song, Wenhui, et al.
Published: (2025)
by: Song, Wenhui, et al.
Published: (2025)
WeatherCity: Urban Scene Reconstruction with Controllable Multi-Weather Transformation
by: Wu, Wenhua, et al.
Published: (2026)
by: Wu, Wenhua, et al.
Published: (2026)
DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles
by: Lin, Chuang, et al.
Published: (2024)
by: Lin, Chuang, et al.
Published: (2024)
Risk-Controllable Multi-View Diffusion for Driving Scenario Generation
by: Lin, Hongyi, et al.
Published: (2026)
by: Lin, Hongyi, et al.
Published: (2026)
MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes
by: Lu, Ruijie, et al.
Published: (2024)
by: Lu, Ruijie, et al.
Published: (2024)
On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection
by: Song, Xiufeng, et al.
Published: (2024)
by: Song, Xiufeng, et al.
Published: (2024)
HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes
by: Soroco, Mauricio, et al.
Published: (2026)
by: Soroco, Mauricio, et al.
Published: (2026)
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
by: Wu, Zehuan, et al.
Published: (2024)
by: Wu, Zehuan, et al.
Published: (2024)
Learning Effective NeRFs and SDFs Representations with 3D Generative Adversarial Networks for 3D Object Generation
by: Yang, Zheyuan, et al.
Published: (2023)
by: Yang, Zheyuan, et al.
Published: (2023)
FreeFix: Boosting 3D Gaussian Splatting via Fine-Tuning-Free Diffusion Models
by: Zhou, Hongyu, et al.
Published: (2026)
by: Zhou, Hongyu, et al.
Published: (2026)
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
by: Ramos, Vasco, et al.
Published: (2024)
by: Ramos, Vasco, et al.
Published: (2024)
MyGo: Consistent and Controllable Multi-View Driving Video Generation with Camera Control
by: Yao, Yining, et al.
Published: (2024)
by: Yao, Yining, et al.
Published: (2024)
MultiWorld: Scalable Multi-Agent Multi-View Video World Models
by: Wu, Haoyu, et al.
Published: (2026)
by: Wu, Haoyu, et al.
Published: (2026)
Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation
by: Qi, Tianhao, et al.
Published: (2025)
by: Qi, Tianhao, et al.
Published: (2025)
CrossModalityDiffusion: Multi-Modal Novel View Synthesis with Unified Intermediate Representation
by: Berian, Alex, et al.
Published: (2025)
by: Berian, Alex, et al.
Published: (2025)
CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
by: Yang, Liudi, et al.
Published: (2025)
by: Yang, Liudi, et al.
Published: (2025)
VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
by: Eskandar, George, et al.
Published: (2026)
by: Eskandar, George, et al.
Published: (2026)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
MultiEgo: A Multi-View Egocentric Video Dataset for 4D Scene Reconstruction
by: Li, Bate, et al.
Published: (2025)
by: Li, Bate, et al.
Published: (2025)
FreeScene: Mixed Graph Diffusion for 3D Scene Synthesis from Free Prompts
by: Bai, Tongyuan, et al.
Published: (2025)
by: Bai, Tongyuan, et al.
Published: (2025)
Similar Items
-
ArmGS: Composite Gaussian Appearance Refinement for Modeling Dynamic Urban Environments
by: Wu, Guile, et al.
Published: (2025) -
Language and Geometry Grounded Sparse Voxel Representations for Holistic Scene Understanding
by: Wu, Guile, et al.
Published: (2026) -
Nighttime Autonomous Driving Scene Reconstruction with Physically-Based Gaussian Splatting
by: Kim, Tae-Kyeong, et al.
Published: (2026) -
TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention
by: Huang, David, et al.
Published: (2026) -
Efficient Depth-Guided Urban View Synthesis
by: Miao, Sheng, et al.
Published: (2024)