ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Xueyi, Zhong, Zuodong, Guo, Yuxin, Liu, Yun-Fu, Su, Zhiguo, Zhang, Qichao, Wang, Junli, Gao, Yinfeng, Zheng, Yupeng, Lin, Qiao, Chen, Huiyong, Zhao, Dongbin
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909799257473024
author Liu, Xueyi
Zhong, Zuodong
Guo, Yuxin
Liu, Yun-Fu
Su, Zhiguo
Zhang, Qichao
Wang, Junli
Gao, Yinfeng
Zheng, Yupeng
Lin, Qiao
Chen, Huiyong
Zhao, Dongbin
author_facet Liu, Xueyi
Zhong, Zuodong
Guo, Yuxin
Liu, Yun-Fu
Su, Zhiguo
Zhang, Qichao
Wang, Junli
Gao, Yinfeng
Zheng, Yupeng
Lin, Qiao
Chen, Huiyong
Zhao, Dongbin
contents Due to the powerful vision-language reasoning and generalization abilities, multimodal large language models (MLLMs) have garnered significant attention in the field of end-to-end (E2E) autonomous driving. However, their application to closed-loop systems remains underexplored, and current MLLM-based methods have not shown clear superiority to mainstream E2E imitation learning approaches. In this work, we propose ReasonPlan, a novel MLLM fine-tuning framework designed for closed-loop driving through holistic reasoning with a self-supervised Next Scene Prediction task and supervised Decision Chain-of-Thought process. This dual mechanism encourages the model to align visual representations with actionable driving context, while promoting interpretable and causally grounded decision making. We curate a planning-oriented decision reasoning dataset, namely PDR, comprising 210k diverse and high-quality samples. Our method outperforms the mainstream E2E imitation learning method by a large margin of 19% L2 and 16.1 driving score on Bench2Drive benchmark. Furthermore, ReasonPlan demonstrates strong zero-shot generalization on unseen DOS benchmark, highlighting its adaptability in handling zero-shot corner cases. Code and dataset will be found in https://github.com/Liuxueyi/ReasonPlan.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20024
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving
Liu, Xueyi
Zhong, Zuodong
Guo, Yuxin
Liu, Yun-Fu
Su, Zhiguo
Zhang, Qichao
Wang, Junli
Gao, Yinfeng
Zheng, Yupeng
Lin, Qiao
Chen, Huiyong
Zhao, Dongbin
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
68T40(Primary), 68T45, 68T50(Secondary)
I.2.9; I.2.10; I.5.1
Due to the powerful vision-language reasoning and generalization abilities, multimodal large language models (MLLMs) have garnered significant attention in the field of end-to-end (E2E) autonomous driving. However, their application to closed-loop systems remains underexplored, and current MLLM-based methods have not shown clear superiority to mainstream E2E imitation learning approaches. In this work, we propose ReasonPlan, a novel MLLM fine-tuning framework designed for closed-loop driving through holistic reasoning with a self-supervised Next Scene Prediction task and supervised Decision Chain-of-Thought process. This dual mechanism encourages the model to align visual representations with actionable driving context, while promoting interpretable and causally grounded decision making. We curate a planning-oriented decision reasoning dataset, namely PDR, comprising 210k diverse and high-quality samples. Our method outperforms the mainstream E2E imitation learning method by a large margin of 19% L2 and 16.1 driving score on Bench2Drive benchmark. Furthermore, ReasonPlan demonstrates strong zero-shot generalization on unseen DOS benchmark, highlighting its adaptability in handling zero-shot corner cases. Code and dataset will be found in https://github.com/Liuxueyi/ReasonPlan.
title ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
68T40(Primary), 68T45, 68T50(Secondary)
I.2.9; I.2.10; I.5.1
url https://arxiv.org/abs/2505.20024