DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Lingjun, Wu, Changjie, Shi, Linzhe, Li, Jiangyang, Liu, Jiaxin, Yang, Lei, Zhang, Hang, Xu, Mu, Wang, Hong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918494598070272
author Zhang, Lingjun
Wu, Changjie
Shi, Linzhe
Li, Jiangyang
Liu, Jiaxin
Yang, Lei
Zhang, Hang
Xu, Mu
Wang, Hong
author_facet Zhang, Lingjun
Wu, Changjie
Shi, Linzhe
Li, Jiangyang
Liu, Jiaxin
Yang, Lei
Zhang, Hang
Xu, Mu
Wang, Hong
contents End-to-end autonomous driving systems are increasingly integrating Vision-Language Model (VLM) architectures, incorporating text reasoning or visual reasoning to enhance the robustness and accuracy of driving decisions. However, the reasoning mechanisms employed in most methods are direct adaptations from general domains, lacking in-depth exploration tailored to autonomous driving scenarios, particularly within visual reasoning modules. In this paper, we propose a driving world model that performs parallel prediction of latent semantic features for consecutive future frames in the bird's-eye-view (BEV) space, thereby enabling long-horizon modeling of future world states. We also introduce an efficient and adaptive text reasoning mechanism that utilizes additional social knowledge and reasoning capabilities to further improve driving performance in challenging long-tail scenarios. We present a novel, efficient, and effective approach that achieves state-of-the-art (SOTA) results on the closed-loop Bench2drive benchmark. Codes are available at: https://github.com/hotdogcheesewhite/DeepSight.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10564
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving
Zhang, Lingjun
Wu, Changjie
Shi, Linzhe
Li, Jiangyang
Liu, Jiaxin
Yang, Lei
Zhang, Hang
Xu, Mu
Wang, Hong
Computer Vision and Pattern Recognition
Robotics
End-to-end autonomous driving systems are increasingly integrating Vision-Language Model (VLM) architectures, incorporating text reasoning or visual reasoning to enhance the robustness and accuracy of driving decisions. However, the reasoning mechanisms employed in most methods are direct adaptations from general domains, lacking in-depth exploration tailored to autonomous driving scenarios, particularly within visual reasoning modules. In this paper, we propose a driving world model that performs parallel prediction of latent semantic features for consecutive future frames in the bird's-eye-view (BEV) space, thereby enabling long-horizon modeling of future world states. We also introduce an efficient and adaptive text reasoning mechanism that utilizes additional social knowledge and reasoning capabilities to further improve driving performance in challenging long-tail scenarios. We present a novel, efficient, and effective approach that achieves state-of-the-art (SOTA) results on the closed-loop Bench2drive benchmark. Codes are available at: https://github.com/hotdogcheesewhite/DeepSight.
title DeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous Driving
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2605.10564