Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Hannan, Wu, Xiaohe, Wang, Shudong, Qin, Xiameng, Zhang, Xinyu, Han, Junyu, Zuo, Wangmeng, Tao, Ji |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generative Inbetweening through Frame-wise Conditions-Driven Video Generation
by: Zhu, Tianyi, et al.
Published: (2024)
by: Zhu, Tianyi, et al.
Published: (2024)
DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes
by: Mei, Jianbiao, et al.
Published: (2024)
by: Mei, Jianbiao, et al.
Published: (2024)
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
by: Zhuo, Dong, et al.
Published: (2026)
by: Zhuo, Dong, et al.
Published: (2026)
SelfHVD: Self-Supervised Handheld Video Deblurring
by: Xu, Honglei, et al.
Published: (2025)
by: Xu, Honglei, et al.
Published: (2025)
Bridging Geometry-Coherent Text-to-3D Generation with Multi-View Diffusion Priors and Gaussian Splatting
by: Yang, Feng, et al.
Published: (2025)
by: Yang, Feng, et al.
Published: (2025)
Evaluation of Text-to-Video Generation Models: A Dynamics Perspective
by: Liao, Mingxiang, et al.
Published: (2024)
by: Liao, Mingxiang, et al.
Published: (2024)
MV-VTON: Multi-View Virtual Try-On with Diffusion Models
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
DiVE: Efficient Multi-View Driving Scenes Generation Based on Video Diffusion Transformer
by: Jiang, Junpeng, et al.
Published: (2025)
by: Jiang, Junpeng, et al.
Published: (2025)
Image Demoiréing Using Dual Camera Fusion on Mobile Phones
by: Mei, Yanting, et al.
Published: (2025)
by: Mei, Yanting, et al.
Published: (2025)
MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer
by: Wu, Guile, et al.
Published: (2025)
by: Wu, Guile, et al.
Published: (2025)
World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving
by: Zhai, Mingliang, et al.
Published: (2024)
by: Zhai, Mingliang, et al.
Published: (2024)
SceneCrafter: Controllable Multi-View Driving Scene Editing
by: Zhu, Zehao, et al.
Published: (2025)
by: Zhu, Zehao, et al.
Published: (2025)
S2AM3D: Scale-controllable Part Segmentation of 3D Point Clouds
by: Su, Han, et al.
Published: (2025)
by: Su, Han, et al.
Published: (2025)
Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models
by: Ding, Xinpeng, et al.
Published: (2024)
by: Ding, Xinpeng, et al.
Published: (2024)
DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
by: Wu, Zehuan, et al.
Published: (2024)
by: Wu, Zehuan, et al.
Published: (2024)
S2Gaussian: Sparse-View Super-Resolution 3D Gaussian Splatting
by: Wan, Yecong, et al.
Published: (2025)
by: Wan, Yecong, et al.
Published: (2025)
SplatWeaver: Learning to Allocate Gaussian Primitives for Generalizable Novel View Synthesis
by: Wan, Yecong, et al.
Published: (2026)
by: Wan, Yecong, et al.
Published: (2026)
MC$^2$: Multi-concept Guidance for Customized Multi-concept Generation
by: Jiang, Jiaxiu, et al.
Published: (2024)
by: Jiang, Jiaxiu, et al.
Published: (2024)
MyGo: Consistent and Controllable Multi-View Driving Video Generation with Camera Control
by: Yao, Yining, et al.
Published: (2024)
by: Yao, Yining, et al.
Published: (2024)
Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
by: Kang, Xueyang, et al.
Published: (2025)
by: Kang, Xueyang, et al.
Published: (2025)
Seeing Across Time and Views: Multi-Temporal Cross-View Learning for Robust Video Person Re-Identification
by: Rashidunnabi, Md, et al.
Published: (2025)
by: Rashidunnabi, Md, et al.
Published: (2025)
FRUC: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views
by: Tao, Yihang, et al.
Published: (2026)
by: Tao, Yihang, et al.
Published: (2026)
MV-S2V: Multi-View Subject-Consistent Video Generation
by: Song, Ziyang, et al.
Published: (2026)
by: Song, Ziyang, et al.
Published: (2026)
FusionTrack: End-to-End Multi-Object Tracking in Arbitrary Multi-View Environment
by: Li, Xiaohe, et al.
Published: (2025)
by: Li, Xiaohe, et al.
Published: (2025)
Reblurring-Guided Single Image Defocus Deblurring: A Learning Framework with Misaligned Training Pairs
by: Ren, Dongwei, et al.
Published: (2024)
by: Ren, Dongwei, et al.
Published: (2024)
VideoMV: Consistent Multi-View Generation Based on Large Video Generative Model
by: Zuo, Qi, et al.
Published: (2024)
by: Zuo, Qi, et al.
Published: (2024)
PostCam: Camera-Controllable Novel-View Video Generation with Query-Shared Cross-Attention
by: Chen, Yipeng, et al.
Published: (2025)
by: Chen, Yipeng, et al.
Published: (2025)
Adaptive Fusion of Single-View and Multi-View Depth for Autonomous Driving
by: Cheng, JunDa, et al.
Published: (2024)
by: Cheng, JunDa, et al.
Published: (2024)
Pseudo-Label Guided Real-World Image De-weathering: A Learning Framework with Imperfect Supervision
by: Xu, Heming, et al.
Published: (2025)
by: Xu, Heming, et al.
Published: (2025)
4Real-Video-V2: Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
by: Feng, Zhiyuan, et al.
Published: (2025)
by: Feng, Zhiyuan, et al.
Published: (2025)
Risk-Controllable Multi-View Diffusion for Driving Scenario Generation
by: Lin, Hongyi, et al.
Published: (2026)
by: Lin, Hongyi, et al.
Published: (2026)
NiteDR: Nighttime Image De-Raining with Cross-View Sensor Cooperative Learning for Dynamic Driving Scenes
by: Shi, Cidan, et al.
Published: (2024)
by: Shi, Cidan, et al.
Published: (2024)
Seeing Beyond Redundancy: Task Complexity's Role in Vision Token Specialization in VLLMs
by: Hannan, Darryl, et al.
Published: (2026)
by: Hannan, Darryl, et al.
Published: (2026)
VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
by: Li, Runjia, et al.
Published: (2025)
by: Li, Runjia, et al.
Published: (2025)
CLIP3D-AD: Extending CLIP for 3D Few-Shot Anomaly Detection with Multi-View Images Generation
by: Zuo, Zuo, et al.
Published: (2024)
by: Zuo, Zuo, et al.
Published: (2024)
ActMVS: Active Scene Reconstruction with Monocular Multi-View Stereo
by: Pu, Guo, et al.
Published: (2026)
by: Pu, Guo, et al.
Published: (2026)
MultiEgo: A Multi-View Egocentric Video Dataset for 4D Scene Reconstruction
by: Li, Bate, et al.
Published: (2025)
by: Li, Bate, et al.
Published: (2025)
HoliGS: Holistic Gaussian Splatting for Embodied View Synthesis
by: Wang, Xiaoyuan, et al.
Published: (2025)
by: Wang, Xiaoyuan, et al.
Published: (2025)
Similar Items
-
Generative Inbetweening through Frame-wise Conditions-Driven Video Generation
by: Zhu, Tianyi, et al.
Published: (2024) -
DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes
by: Mei, Jianbiao, et al.
Published: (2024) -
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
by: Zhuo, Dong, et al.
Published: (2026) -
SelfHVD: Self-Supervised Handheld Video Deblurring
by: Xu, Honglei, et al.
Published: (2025) -
Bridging Geometry-Coherent Text-to-3D Generation with Multi-View Diffusion Priors and Gaussian Splatting
by: Yang, Feng, et al.
Published: (2025)