HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Xin, Liang, Dingkang, Tu, Sifan, Chen, Xiwu, Ding, Yikang, Zhang, Dingyuan, Tan, Feiyang, Zhao, Hengshuang, Bai, Xiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
von: Zhou, Xin, et al.
Veröffentlicht: (2026)
von: Zhou, Xin, et al.
Veröffentlicht: (2026)
A Unified Framework for 3D Scene Understanding
von: Xu, Wei, et al.
Veröffentlicht: (2024)
von: Xu, Wei, et al.
Veröffentlicht: (2024)
UniFuture: A 4D Driving World Model for Future Generation and Perception
von: Liang, Dingkang, et al.
Veröffentlicht: (2025)
von: Liang, Dingkang, et al.
Veröffentlicht: (2025)
Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey
von: Tu, Sifan, et al.
Veröffentlicht: (2025)
von: Tu, Sifan, et al.
Veröffentlicht: (2025)
UniScene: Unified Occupancy-centric Driving Scene Generation
von: Li, Bohan, et al.
Veröffentlicht: (2024)
von: Li, Bohan, et al.
Veröffentlicht: (2024)
DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene Generation
von: Guo, Jiazhe, et al.
Veröffentlicht: (2025)
von: Guo, Jiazhe, et al.
Veröffentlicht: (2025)
A Unified Image-Dense Annotation Generation Model for Underwater Scenes
von: Lin, Hongkai, et al.
Veröffentlicht: (2025)
von: Lin, Hongkai, et al.
Veröffentlicht: (2025)
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
von: Chen, Kaijin, et al.
Veröffentlicht: (2026)
von: Chen, Kaijin, et al.
Veröffentlicht: (2026)
PointTPA: Dynamic Network Parameter Adaptation for 3D Scene Understanding
von: Liu, Siyuan, et al.
Veröffentlicht: (2026)
von: Liu, Siyuan, et al.
Veröffentlicht: (2026)
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
von: Zhao, Zongchuang, et al.
Veröffentlicht: (2025)
von: Zhao, Zongchuang, et al.
Veröffentlicht: (2025)
Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding
von: Wu, Xianjin, et al.
Veröffentlicht: (2026)
von: Wu, Xianjin, et al.
Veröffentlicht: (2026)
More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models
von: Lin, Hongkai, et al.
Veröffentlicht: (2025)
von: Lin, Hongkai, et al.
Veröffentlicht: (2025)
AVS-Net: Point Sampling with Adaptive Voxel Size for 3D Scene Understanding
von: Yang, Hongcheng, et al.
Veröffentlicht: (2024)
von: Yang, Hongcheng, et al.
Veröffentlicht: (2024)
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
von: Sun, Zhengyang, et al.
Veröffentlicht: (2026)
von: Sun, Zhengyang, et al.
Veröffentlicht: (2026)
Make Your ViT-based Multi-view 3D Detectors Faster via Token Compression
von: Zhang, Dingyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Dingyuan, et al.
Veröffentlicht: (2024)
Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation
von: Gui, Xingtai, et al.
Veröffentlicht: (2026)
von: Gui, Xingtai, et al.
Veröffentlicht: (2026)
DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning
von: Liu, Zhe, et al.
Veröffentlicht: (2025)
von: Liu, Zhe, et al.
Veröffentlicht: (2025)
SAM3D: Zero-Shot 3D Object Detection via Segment Anything Model
von: Zhang, Dingyuan, et al.
Veröffentlicht: (2023)
von: Zhang, Dingyuan, et al.
Veröffentlicht: (2023)
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
von: Fu, Haoyu, et al.
Veröffentlicht: (2025)
von: Fu, Haoyu, et al.
Veröffentlicht: (2025)
NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding
von: Xu, Wei, et al.
Veröffentlicht: (2025)
von: Xu, Wei, et al.
Veröffentlicht: (2025)
GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation
von: Yang, Zhenya, et al.
Veröffentlicht: (2025)
von: Yang, Zhenya, et al.
Veröffentlicht: (2025)
UniLION: Towards Unified Autonomous Driving Model with Linear Group RNNs
von: Liu, Zhe, et al.
Veröffentlicht: (2025)
von: Liu, Zhe, et al.
Veröffentlicht: (2025)
You Only Look Bottom-Up for Monocular 3D Object Detection
von: Xiong, Kaixin, et al.
Veröffentlicht: (2024)
von: Xiong, Kaixin, et al.
Veröffentlicht: (2024)
GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
von: Deng, Tianchen, et al.
Veröffentlicht: (2025)
PlayerOne: Egocentric World Simulator
von: Tu, Yuanpeng, et al.
Veröffentlicht: (2025)
von: Tu, Yuanpeng, et al.
Veröffentlicht: (2025)
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
von: Zhuo, Dong, et al.
Veröffentlicht: (2026)
von: Zhuo, Dong, et al.
Veröffentlicht: (2026)
DriveWorld: 4D Pre-trained Scene Understanding via World Models for Autonomous Driving
von: Min, Chen, et al.
Veröffentlicht: (2024)
von: Min, Chen, et al.
Veröffentlicht: (2024)
UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving
von: Xiong, Zhexiao, et al.
Veröffentlicht: (2026)
von: Xiong, Zhexiao, et al.
Veröffentlicht: (2026)
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
von: Qi, Zhangyang, et al.
Veröffentlicht: (2025)
von: Qi, Zhangyang, et al.
Veröffentlicht: (2025)
CLAIM: Camera-LiDAR Alignment with Intensity and Monodepth
von: Zhang, Zhuo, et al.
Veröffentlicht: (2025)
von: Zhang, Zhuo, et al.
Veröffentlicht: (2025)
ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving
von: Huang, Minqing, et al.
Veröffentlicht: (2026)
von: Huang, Minqing, et al.
Veröffentlicht: (2026)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
von: jia, Feiyang, et al.
Veröffentlicht: (2026)
von: jia, Feiyang, et al.
Veröffentlicht: (2026)
MuDG: Taming Multi-modal Diffusion with Gaussian Splatting for Urban Scene Reconstruction
von: Zou, Yingshuang, et al.
Veröffentlicht: (2025)
von: Zou, Yingshuang, et al.
Veröffentlicht: (2025)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
Liquid: Language Models are Scalable and Unified Multi-modal Generators
von: Wu, Junfeng, et al.
Veröffentlicht: (2024)
von: Wu, Junfeng, et al.
Veröffentlicht: (2024)
Unified 3D Scene Understanding Through Physical World Modeling
von: Lee, Wanhee, et al.
Veröffentlicht: (2026)
von: Lee, Wanhee, et al.
Veröffentlicht: (2026)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
von: Ospanov, Azim, et al.
Veröffentlicht: (2025)
von: Ospanov, Azim, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
von: Zhou, Xin, et al.
Veröffentlicht: (2026) -
A Unified Framework for 3D Scene Understanding
von: Xu, Wei, et al.
Veröffentlicht: (2024) -
UniFuture: A 4D Driving World Model for Future Generation and Perception
von: Liang, Dingkang, et al.
Veröffentlicht: (2025) -
Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching
von: Zhou, Xin, et al.
Veröffentlicht: (2025) -
The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey
von: Tu, Sifan, et al.
Veröffentlicht: (2025)