AD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yan, Tianyi, Tang, Tao, Gui, Xingtai, Li, Yongkang, Zhesng, Jiasen, Huang, Weiyao, Kong, Lingdong, Han, Wencheng, Zhou, Xia, Zhang, Xueyang, Zhan, Yifei, Zhan, Kun, Xu, Cheng-zhong, Shen, Jianbing
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917330726944768
author Yan, Tianyi
Tang, Tao
Gui, Xingtai
Li, Yongkang
Zhesng, Jiasen
Huang, Weiyao
Kong, Lingdong
Han, Wencheng
Zhou, Xia
Zhang, Xueyang
Zhan, Yifei
Zhan, Kun
Xu, Cheng-zhong
Shen, Jianbing
author_facet Yan, Tianyi
Tang, Tao
Gui, Xingtai
Li, Yongkang
Zhesng, Jiasen
Huang, Weiyao
Kong, Lingdong
Han, Wencheng
Zhou, Xia
Zhang, Xueyang
Zhan, Yifei
Zhan, Kun
Xu, Cheng-zhong
Shen, Jianbing
contents End-to-end models for autonomous driving hold the promise of learning complex behaviors directly from sensor data, but face critical challenges in safety and handling long-tail events. Reinforcement Learning (RL) offers a promising path to overcome these limitations, yet its success in autonomous driving has been elusive. We identify a fundamental flaw hindering this progress: a deep seated optimistic bias in the world models used for RL. To address this, we introduce a framework for post-training policy refinement built around an Impartial World Model. Our primary contribution is to teach this model to be honest about danger. We achieve this with a novel data synthesis pipeline, Counterfactual Synthesis, which systematically generates a rich curriculum of plausible collisions and off-road events. This transforms the model from a passive scene completer into a veridical forecaster that remains faithful to the causal link between actions and outcomes. We then integrate this Impartial World Model into our closed-loop RL framework, where it serves as an internal critic. During refinement, the agent queries the critic to ``dream" of the outcomes for candidate actions. We demonstrate through extensive experiments, including on a new Risk Foreseeing Benchmark, that our model significantly outperforms baselines in predicting failures. Consequently, when used as a critic, it enables a substantial reduction in safety violations in challenging simulations, proving that teaching a model to dream of danger is a critical step towards building truly safe and intelligent autonomous agents.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20325
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models
Yan, Tianyi
Tang, Tao
Gui, Xingtai
Li, Yongkang
Zhesng, Jiasen
Huang, Weiyao
Kong, Lingdong
Han, Wencheng
Zhou, Xia
Zhang, Xueyang
Zhan, Yifei
Zhan, Kun
Xu, Cheng-zhong
Shen, Jianbing
Computer Vision and Pattern Recognition
End-to-end models for autonomous driving hold the promise of learning complex behaviors directly from sensor data, but face critical challenges in safety and handling long-tail events. Reinforcement Learning (RL) offers a promising path to overcome these limitations, yet its success in autonomous driving has been elusive. We identify a fundamental flaw hindering this progress: a deep seated optimistic bias in the world models used for RL. To address this, we introduce a framework for post-training policy refinement built around an Impartial World Model. Our primary contribution is to teach this model to be honest about danger. We achieve this with a novel data synthesis pipeline, Counterfactual Synthesis, which systematically generates a rich curriculum of plausible collisions and off-road events. This transforms the model from a passive scene completer into a veridical forecaster that remains faithful to the causal link between actions and outcomes. We then integrate this Impartial World Model into our closed-loop RL framework, where it serves as an internal critic. During refinement, the agent queries the critic to ``dream" of the outcomes for candidate actions. We demonstrate through extensive experiments, including on a new Risk Foreseeing Benchmark, that our model significantly outperforms baselines in predicting failures. Consequently, when used as a critic, it enables a substantial reduction in safety violations in challenging simulations, proving that teaching a model to dream of danger is a critical step towards building truly safe and intelligent autonomous agents.
title AD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.20325