RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Hao, Chen, Shaoyu, Jiang, Bo, Liao, Bencheng, Shi, Yiang, Guo, Xiaoyang, Pu, Yuechuan, Yin, Haoran, Li, Xiangyu, Zhang, Xinbang, Zhang, Ying, Liu, Wenyu, Zhang, Qian, Wang, Xinggang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909859311517696
author Gao, Hao
Chen, Shaoyu
Jiang, Bo
Liao, Bencheng
Shi, Yiang
Guo, Xiaoyang
Pu, Yuechuan
Yin, Haoran
Li, Xiangyu
Zhang, Xinbang
Zhang, Ying
Liu, Wenyu
Zhang, Qian
Wang, Xinggang
author_facet Gao, Hao
Chen, Shaoyu
Jiang, Bo
Liao, Bencheng
Shi, Yiang
Guo, Xiaoyang
Pu, Yuechuan
Yin, Haoran
Li, Xiangyu
Zhang, Xinbang
Zhang, Ying
Liu, Wenyu
Zhang, Qian
Wang, Xinggang
contents Existing end-to-end autonomous driving (AD) algorithms typically follow the Imitation Learning (IL) paradigm, which faces challenges such as causal confusion and an open-loop gap. In this work, we propose RAD, a 3DGS-based closed-loop Reinforcement Learning (RL) framework for end-to-end Autonomous Driving. By leveraging 3DGS techniques, we construct a photorealistic digital replica of the real physical world, enabling the AD policy to extensively explore the state space and learn to handle out-of-distribution scenarios through large-scale trial and error. To enhance safety, we design specialized rewards to guide the policy in effectively responding to safety-critical events and understanding real-world causal relationships. To better align with human driving behavior, we incorporate IL into RL training as a regularization term. We introduce a closed-loop evaluation benchmark consisting of diverse, previously unseen 3DGS environments. Compared to IL-based methods, RAD achieves stronger performance in most closed-loop metrics, particularly exhibiting a 3x lower collision rate. Abundant closed-loop results are presented in the supplementary material. Code is available at https://github.com/hustvl/RAD for facilitating future research.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13144
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning
Gao, Hao
Chen, Shaoyu
Jiang, Bo
Liao, Bencheng
Shi, Yiang
Guo, Xiaoyang
Pu, Yuechuan
Yin, Haoran
Li, Xiangyu
Zhang, Xinbang
Zhang, Ying
Liu, Wenyu
Zhang, Qian
Wang, Xinggang
Computer Vision and Pattern Recognition
Robotics
Existing end-to-end autonomous driving (AD) algorithms typically follow the Imitation Learning (IL) paradigm, which faces challenges such as causal confusion and an open-loop gap. In this work, we propose RAD, a 3DGS-based closed-loop Reinforcement Learning (RL) framework for end-to-end Autonomous Driving. By leveraging 3DGS techniques, we construct a photorealistic digital replica of the real physical world, enabling the AD policy to extensively explore the state space and learn to handle out-of-distribution scenarios through large-scale trial and error. To enhance safety, we design specialized rewards to guide the policy in effectively responding to safety-critical events and understanding real-world causal relationships. To better align with human driving behavior, we incorporate IL into RL training as a regularization term. We introduce a closed-loop evaluation benchmark consisting of diverse, previously unseen 3DGS environments. Compared to IL-based methods, RAD achieves stronger performance in most closed-loop metrics, particularly exhibiting a 3x lower collision rate. Abundant closed-loop results are presented in the supplementary material. Code is available at https://github.com/hustvl/RAD for facilitating future research.
title RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2502.13144