Generating Multimodal Driving Scenes via Next-Scene Prediction
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Yanhao, Zhang, Haoyang, Lin, Tianwei, Huang, Lichao, Luo, Shujie, Wu, Rui, Qiu, Congpei, Ke, Wei, Zhang, Tong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
por: Wu, Yanhao, et al.
Publicado: (2026)
por: Wu, Yanhao, et al.
Publicado: (2026)
Refining CLIP's Spatial Awareness: A Visual-Centric Perspective
por: Qiu, Congpei, et al.
Publicado: (2025)
por: Qiu, Congpei, et al.
Publicado: (2025)
UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register
por: Qiu, Congpei, et al.
Publicado: (2026)
por: Qiu, Congpei, et al.
Publicado: (2026)
Mitigating Object Dependencies: Improving Point Cloud Self-Supervised Learning through Object Exchange
por: Wu, Yanhao, et al.
Publicado: (2024)
por: Wu, Yanhao, et al.
Publicado: (2024)
Pair2Scene: Learning Local Object Relations for Procedural Scene Generation
por: Ran, Xingjian, et al.
Publicado: (2026)
por: Ran, Xingjian, et al.
Publicado: (2026)
OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving
por: Zheng, Lianqing, et al.
Publicado: (2024)
por: Zheng, Lianqing, et al.
Publicado: (2024)
SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation
por: Sun, Wenchao, et al.
Publicado: (2024)
por: Sun, Wenchao, et al.
Publicado: (2024)
Physics-Aware 3D Gaussian Editing for Driving Scene Generation
por: Zhou, Feng, et al.
Publicado: (2026)
por: Zhou, Feng, et al.
Publicado: (2026)
UniScene: Unified Occupancy-centric Driving Scene Generation
por: Li, Bohan, et al.
Publicado: (2024)
por: Li, Bohan, et al.
Publicado: (2024)
ReinDriveGen: Reinforcement Post-Training for Out-of-Distribution Driving Scene Generation
por: Zhang, Hao, et al.
Publicado: (2026)
por: Zhang, Hao, et al.
Publicado: (2026)
SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent
por: Yang, Yandan, et al.
Publicado: (2025)
por: Yang, Yandan, et al.
Publicado: (2025)
What Happens Next? Next Scene Prediction with a Unified Video Model
por: Li, Xinjie, et al.
Publicado: (2025)
por: Li, Xinjie, et al.
Publicado: (2025)
GA-Drive: Geometry-Appearance Decoupled Modeling for Free-viewpoint Driving Scene Generation
por: Zhang, Hao, et al.
Publicado: (2026)
por: Zhang, Hao, et al.
Publicado: (2026)
DHGS: Decoupled Hybrid Gaussian Splatting for Driving Scene
por: Shi, Xi, et al.
Publicado: (2024)
por: Shi, Xi, et al.
Publicado: (2024)
BEVWorld: A Multimodal World Simulator for Autonomous Driving via Scene-Level BEV Latents
por: Zhang, Yumeng, et al.
Publicado: (2024)
por: Zhang, Yumeng, et al.
Publicado: (2024)
Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
por: Guo, Xiangyu, et al.
Publicado: (2025)
por: Guo, Xiangyu, et al.
Publicado: (2025)
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes
por: Wang, Chuhan, et al.
Publicado: (2026)
por: Wang, Chuhan, et al.
Publicado: (2026)
3D Scene Change Modeling With Consistent Multi-View Aggregation
por: Zhou, Zirui, et al.
Publicado: (2025)
por: Zhou, Zirui, et al.
Publicado: (2025)
Predicting 3D representations for Dynamic Scenes
por: Qi, Di, et al.
Publicado: (2025)
por: Qi, Di, et al.
Publicado: (2025)
InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving
por: Song, Ruiqi, et al.
Publicado: (2025)
por: Song, Ruiqi, et al.
Publicado: (2025)
Masked Next-Scale Prediction for Self-supervised Scene Text Recognition
por: Chen, Zhuohao, et al.
Publicado: (2026)
por: Chen, Zhuohao, et al.
Publicado: (2026)
OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving
por: Liu, Pei, et al.
Publicado: (2025)
por: Liu, Pei, et al.
Publicado: (2025)
AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond
por: Zhang, Haiming, et al.
Publicado: (2026)
por: Zhang, Haiming, et al.
Publicado: (2026)
SGG-R$^{\rm 3}$: From Next-Token Prediction to End-to-End Unbiased Scene Graph Generation
por: Feng, Jiaye, et al.
Publicado: (2026)
por: Feng, Jiaye, et al.
Publicado: (2026)
SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Driving
por: Li, Jingyu, et al.
Publicado: (2026)
por: Li, Jingyu, et al.
Publicado: (2026)
DreamScene: 3D Gaussian-based Text-to-3D Scene Generation via Formation Pattern Sampling
por: Li, Haoran, et al.
Publicado: (2024)
por: Li, Haoran, et al.
Publicado: (2024)
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
por: Wu, Zehuan, et al.
Publicado: (2024)
por: Wu, Zehuan, et al.
Publicado: (2024)
Expanding Scene Graph Boundaries: Fully Open-vocabulary Scene Graph Generation via Visual-Concept Alignment and Retention
por: Chen, Zuyao, et al.
Publicado: (2023)
por: Chen, Zuyao, et al.
Publicado: (2023)
SceneStreamer: Continuous Scenario Generation as Next Token Group Prediction
por: Peng, Zhenghao, et al.
Publicado: (2025)
por: Peng, Zhenghao, et al.
Publicado: (2025)
GS-RoadPatching: Inpainting Gaussians via 3D Searching and Placing for Driving Scenes
por: Chen, Guo, et al.
Publicado: (2025)
por: Chen, Guo, et al.
Publicado: (2025)
MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization
por: Xiao, Zhendong, et al.
Publicado: (2025)
por: Xiao, Zhendong, et al.
Publicado: (2025)
Unsupervised Collaborative Domain Adaptation for Driving Scene Parsing
por: Fan, Jiahe, et al.
Publicado: (2026)
por: Fan, Jiahe, et al.
Publicado: (2026)
SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments
por: Cao, Yue, et al.
Publicado: (2024)
por: Cao, Yue, et al.
Publicado: (2024)
4D Driving Scene Generation With Stereo Forcing
por: Lu, Hao, et al.
Publicado: (2025)
por: Lu, Hao, et al.
Publicado: (2025)
DriveSplat: Unified Neural Gaussian Reconstruction for Dynamic Driving Scenes
por: Wang, Cong, et al.
Publicado: (2025)
por: Wang, Cong, et al.
Publicado: (2025)
SceneCrafter: Controllable Multi-View Driving Scene Editing
por: Zhu, Zehao, et al.
Publicado: (2025)
por: Zhu, Zehao, et al.
Publicado: (2025)
The Scene Language: Representing Scenes with Programs, Words, and Embeddings
por: Zhang, Yunzhi, et al.
Publicado: (2024)
por: Zhang, Yunzhi, et al.
Publicado: (2024)
FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow
por: Yang, Zhifei, et al.
Publicado: (2026)
por: Yang, Zhifei, et al.
Publicado: (2026)
Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention
por: Lu, Hannan, et al.
Publicado: (2024)
por: Lu, Hannan, et al.
Publicado: (2024)
Active Learning from Scene Embeddings for End-to-End Autonomous Driving
por: Jiang, Wenhao, et al.
Publicado: (2025)
por: Jiang, Wenhao, et al.
Publicado: (2025)
Ejemplares similares
-
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
por: Wu, Yanhao, et al.
Publicado: (2026) -
Refining CLIP's Spatial Awareness: A Visual-Centric Perspective
por: Qiu, Congpei, et al.
Publicado: (2025) -
UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register
por: Qiu, Congpei, et al.
Publicado: (2026) -
Mitigating Object Dependencies: Improving Point Cloud Self-Supervised Learning through Object Exchange
por: Wu, Yanhao, et al.
Publicado: (2024) -
Pair2Scene: Learning Local Object Relations for Procedural Scene Generation
por: Ran, Xingjian, et al.
Publicado: (2026)