DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918494598070272 |
|---|---|
| author | Zhang, Lingjun Wu, Changjie Shi, Linzhe Li, Jiangyang Liu, Jiaxin Yang, Lei Zhang, Hang Xu, Mu Wang, Hong |
| author_facet | Zhang, Lingjun Wu, Changjie Shi, Linzhe Li, Jiangyang Liu, Jiaxin Yang, Lei Zhang, Hang Xu, Mu Wang, Hong |
| contents | End-to-end autonomous driving systems are increasingly integrating Vision-Language Model (VLM) architectures, incorporating text reasoning or visual reasoning to enhance the robustness and accuracy of driving decisions. However, the reasoning mechanisms employed in most methods are direct adaptations from general domains, lacking in-depth exploration tailored to autonomous driving scenarios, particularly within visual reasoning modules. In this paper, we propose a driving world model that performs parallel prediction of latent semantic features for consecutive future frames in the bird's-eye-view (BEV) space, thereby enabling long-horizon modeling of future world states. We also introduce an efficient and adaptive text reasoning mechanism that utilizes additional social knowledge and reasoning capabilities to further improve driving performance in challenging long-tail scenarios. We present a novel, efficient, and effective approach that achieves state-of-the-art (SOTA) results on the closed-loop Bench2drive benchmark. Codes are available at: https://github.com/hotdogcheesewhite/DeepSight. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_10564 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving Zhang, Lingjun Wu, Changjie Shi, Linzhe Li, Jiangyang Liu, Jiaxin Yang, Lei Zhang, Hang Xu, Mu Wang, Hong Computer Vision and Pattern Recognition Robotics End-to-end autonomous driving systems are increasingly integrating Vision-Language Model (VLM) architectures, incorporating text reasoning or visual reasoning to enhance the robustness and accuracy of driving decisions. However, the reasoning mechanisms employed in most methods are direct adaptations from general domains, lacking in-depth exploration tailored to autonomous driving scenarios, particularly within visual reasoning modules. In this paper, we propose a driving world model that performs parallel prediction of latent semantic features for consecutive future frames in the bird's-eye-view (BEV) space, thereby enabling long-horizon modeling of future world states. We also introduce an efficient and adaptive text reasoning mechanism that utilizes additional social knowledge and reasoning capabilities to further improve driving performance in challenging long-tail scenarios. We present a novel, efficient, and effective approach that achieves state-of-the-art (SOTA) results on the closed-loop Bench2drive benchmark. Codes are available at: https://github.com/hotdogcheesewhite/DeepSight. |
| title | DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving |
| topic | Computer Vision and Pattern Recognition Robotics |
| url | https://arxiv.org/abs/2605.10564 |