See Tomorrow, Act Today: Foresight-Driven Autonomous Driving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Bozhou, Song, Nan, Wang, Yuang, Deng, Jiankang, Zhu, Xiatian, Zhang, Li
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918489192660992
author Zhang, Bozhou
Song, Nan
Wang, Yuang
Deng, Jiankang
Zhu, Xiatian
Zhang, Li
author_facet Zhang, Bozhou
Song, Nan
Wang, Yuang
Deng, Jiankang
Zhu, Xiatian
Zhang, Li
contents Current end-to-end autonomous driving planners are fundamentally reactive: they condition on historical and present observations to predict future actions. We argue that autonomous agents should instead imagine future scenes before deciding, just as human drivers mentally simulate ``what will happen next" before acting. We introduce ForeSight, a foundation world model centric planning framework that reframes autonomous driving as anticipatory decision-making. Rather than treating world models as auxiliary components, ForeSight makes future scene imagination the primary driver of action prediction. Our approach operates in two stages: (1) generating plausible future visual worlds via a pretrained world model, and (2) planning actions conditioned on these imagined futures. This paradigm shift from ``what should I do now?" to ``what will happen, and how should I respond?" enables genuinely anticipatory rather than reactive planning. By grounding decisions in anticipated contexts rather than present observations alone, ForeSight navigates dynamic, interactive scenarios more effectively. Extensive experiments on NAVSIM and nuScenes demonstrate that explicit future imagination significantly outperforms previous state-of-the-art alternatives, validating our foresight-driven approach.
format Preprint
id arxiv_https___arxiv_org_abs_2605_07195
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle See Tomorrow, Act Today: Foresight-Driven Autonomous Driving
Zhang, Bozhou
Song, Nan
Wang, Yuang
Deng, Jiankang
Zhu, Xiatian
Zhang, Li
Computer Vision and Pattern Recognition
Current end-to-end autonomous driving planners are fundamentally reactive: they condition on historical and present observations to predict future actions. We argue that autonomous agents should instead imagine future scenes before deciding, just as human drivers mentally simulate ``what will happen next" before acting. We introduce ForeSight, a foundation world model centric planning framework that reframes autonomous driving as anticipatory decision-making. Rather than treating world models as auxiliary components, ForeSight makes future scene imagination the primary driver of action prediction. Our approach operates in two stages: (1) generating plausible future visual worlds via a pretrained world model, and (2) planning actions conditioned on these imagined futures. This paradigm shift from ``what should I do now?" to ``what will happen, and how should I respond?" enables genuinely anticipatory rather than reactive planning. By grounding decisions in anticipated contexts rather than present observations alone, ForeSight navigates dynamic, interactive scenarios more effectively. Extensive experiments on NAVSIM and nuScenes demonstrate that explicit future imagination significantly outperforms previous state-of-the-art alternatives, validating our foresight-driven approach.
title See Tomorrow, Act Today: Foresight-Driven Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.07195