Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Xiangyu, Wu, Zhanqian, Xiong, Kaixin, Xu, Ziyang, Zhou, Lijun, Xu, Gangwei, Xu, Shaoqing, Sun, Haiyang, Wang, Bing, Chen, Guang, Ye, Hangjun, Liu, Wenyu, Wang, Xinggang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
by: Li, Yongkang, et al.
Published: (2025)
by: Li, Yongkang, et al.
Published: (2025)
DriveLaW:Unifying Planning and Video Generation in a Latent Driving World
by: Xia, Tianze, et al.
Published: (2025)
by: Xia, Tianze, et al.
Published: (2025)
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
by: Li, Yongkang, et al.
Published: (2026)
by: Li, Yongkang, et al.
Published: (2026)
WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving
by: Zhu, Ziyue, et al.
Published: (2025)
by: Zhu, Ziyue, et al.
Published: (2025)
Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks
by: Zeng, Kai, et al.
Published: (2025)
by: Zeng, Kai, et al.
Published: (2025)
STP4D: Spatio-Temporal-Prompt Consistent Modeling for Text-to-4D Gaussian Splatting
by: Deng, Yunze, et al.
Published: (2025)
by: Deng, Yunze, et al.
Published: (2025)
Toward Physically Consistent Driving Video World Models under Challenging Trajectories
by: Zhou, Jiawei, et al.
Published: (2026)
by: Zhou, Jiawei, et al.
Published: (2026)
Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual Odometry
by: Zhang, Zhaoxing, et al.
Published: (2024)
by: Zhang, Zhaoxing, et al.
Published: (2024)
LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
by: Luo, Yuechen, et al.
Published: (2026)
by: Luo, Yuechen, et al.
Published: (2026)
Mirage: One-Step Video Diffusion for Photorealistic and Coherent Asset Editing in Driving Scenes
by: Wang, Shuyun, et al.
Published: (2025)
by: Wang, Shuyun, et al.
Published: (2025)
GaraMoSt: Parallel Multi-Granularity Motion and Structural Modeling for Efficient Multi-Frame Interpolation in DSA Images
by: Xu, Ziyang, et al.
Published: (2024)
by: Xu, Ziyang, et al.
Published: (2024)
PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations
by: Chi, Cheng, et al.
Published: (2026)
by: Chi, Cheng, et al.
Published: (2026)
Pixel-Perfect Visual Geometry Estimation
by: Xu, Gangwei, et al.
Published: (2026)
by: Xu, Gangwei, et al.
Published: (2026)
UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene Modeling
by: Tan, Kaiyuan, et al.
Published: (2026)
by: Tan, Kaiyuan, et al.
Published: (2026)
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
by: Hou, Zhiyi, et al.
Published: (2025)
by: Hou, Zhiyi, et al.
Published: (2025)
CoGen: 3D Consistent Video Generation via Adaptive Conditioning for Autonomous Driving
by: Ji, Yishen, et al.
Published: (2025)
by: Ji, Yishen, et al.
Published: (2025)
PixelHacker: Image Inpainting with Structural and Semantic Consistency
by: Xu, Ziyang, et al.
Published: (2025)
by: Xu, Ziyang, et al.
Published: (2025)
Deflickering Vision-Based Occupancy Networks through Lightweight Spatio-Temporal Correlation
by: Yu, Fengcheng, et al.
Published: (2025)
by: Yu, Fengcheng, et al.
Published: (2025)
MoSt-DSA: Modeling Motion and Structural Interactions for Direct Multi-Frame Interpolation in DSA Images
by: Xu, Ziyang, et al.
Published: (2024)
by: Xu, Ziyang, et al.
Published: (2024)
XS-VID: An Extremely Small Video Object Detection Dataset
by: Guo, Jiahao, et al.
Published: (2024)
by: Guo, Jiahao, et al.
Published: (2024)
DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images
by: Chen, Xiaoxue, et al.
Published: (2025)
by: Chen, Xiaoxue, et al.
Published: (2025)
GaitGS: Temporal Feature Learning in Granularity and Span Dimension for Gait Recognition
by: Xiong, Haijun, et al.
Published: (2023)
by: Xiong, Haijun, et al.
Published: (2023)
DriveFix: Spatio-Temporally Coherent Driving Scene Restoration
by: Si, Heyu, et al.
Published: (2026)
by: Si, Heyu, et al.
Published: (2026)
BAT: Learning Event-based Optical Flow with Bidirectional Adaptive Temporal Correlation
by: Xu, Gangwei, et al.
Published: (2025)
by: Xu, Gangwei, et al.
Published: (2025)
FlowMamba: Learning Point Cloud Scene Flow with Global Motion Propagation
by: Lin, Min, et al.
Published: (2024)
by: Lin, Min, et al.
Published: (2024)
AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
by: Jiang, Bo, et al.
Published: (2025)
by: Jiang, Bo, et al.
Published: (2025)
Non-Cross Diffusion for Semantic Consistency
by: Zheng, Ziyang, et al.
Published: (2023)
by: Zheng, Ziyang, et al.
Published: (2023)
LiDAR Prompted Spatio-Temporal Multi-View Stereo for Autonomous Driving
by: Sun, Qihao, et al.
Published: (2026)
by: Sun, Qihao, et al.
Published: (2026)
ViSE: A Systematic Approach to Vision-Only Street-View Extrapolation
by: Tan, Kaiyuan, et al.
Published: (2025)
by: Tan, Kaiyuan, et al.
Published: (2025)
PCSTracker: Long-Term Scene Flow Estimation for Point Cloud Sequences
by: Lin, Min, et al.
Published: (2026)
by: Lin, Min, et al.
Published: (2026)
Fast High Dynamic Range Radiance Fields for Dynamic Scenes
by: Wu, Guanjun, et al.
Published: (2024)
by: Wu, Guanjun, et al.
Published: (2024)
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
by: Cheng, Tianheng, et al.
Published: (2026)
by: Cheng, Tianheng, et al.
Published: (2026)
Causality-inspired Discriminative Feature Learning in Triple Domains for Gait Recognition
by: Xiong, Haijun, et al.
Published: (2024)
by: Xiong, Haijun, et al.
Published: (2024)
Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
by: Xu, Gangwei, et al.
Published: (2025)
by: Xu, Gangwei, et al.
Published: (2025)
Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification
by: Wang, Xiao, et al.
Published: (2026)
by: Wang, Xiao, et al.
Published: (2026)
BEVWorld: A Multimodal World Simulator for Autonomous Driving via Scene-Level BEV Latents
by: Zhang, Yumeng, et al.
Published: (2024)
by: Zhang, Yumeng, et al.
Published: (2024)
MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
by: Zou, Jialv, et al.
Published: (2024)
by: Zou, Jialv, et al.
Published: (2024)
Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft
by: Huang, Junchao, et al.
Published: (2025)
by: Huang, Junchao, et al.
Published: (2025)
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
by: Mei, Xiaodong, et al.
Published: (2026)
by: Mei, Xiaodong, et al.
Published: (2026)
Gait Recognition via Collaborating Discriminative and Generative Diffusion Models
by: Xiong, Haijun, et al.
Published: (2025)
by: Xiong, Haijun, et al.
Published: (2025)
Similar Items
-
ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
by: Li, Yongkang, et al.
Published: (2025) -
DriveLaW:Unifying Planning and Video Generation in a Latent Driving World
by: Xia, Tianze, et al.
Published: (2025) -
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
by: Li, Yongkang, et al.
Published: (2026) -
WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving
by: Zhu, Ziyue, et al.
Published: (2025) -
Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks
by: Zeng, Kai, et al.
Published: (2025)