Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-lane Scenarios

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Feihong, Zhan, Guojian, Shuai, Bin, Zhang, Tianyi, Duan, Jingliang, Li, Shengbo Eben
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909617414471680
author Zhang, Feihong
Zhan, Guojian
Shuai, Bin
Zhang, Tianyi
Duan, Jingliang
Li, Shengbo Eben
author_facet Zhang, Feihong
Zhan, Guojian
Shuai, Bin
Zhang, Tianyi
Duan, Jingliang
Li, Shengbo Eben
contents Reinforcement learning (RL), known for its self-evolution capability, offers a promising approach to training high-level autonomous driving systems. However, handling constraints remains a significant challenge for existing RL algorithms, particularly in real-world applications. In this paper, we propose a new safety-oriented training technique called harmonic policy iteration (HPI). At each RL iteration, it first calculates two policy gradients associated with efficient driving and safety constraints, respectively. Then, a harmonic gradient is derived for policy updating, minimizing conflicts between the two gradients and consequently enabling a more balanced and stable training process. Furthermore, we adopt the state-of-the-art DSAC algorithm as the backbone and integrate it with our HPI to develop a new safe RL algorithm, DSAC-H. Extensive simulations in multi-lane scenarios demonstrate that DSAC-H achieves efficient driving performance with near-zero safety constraint violations.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13532
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-lane Scenarios
Zhang, Feihong
Zhan, Guojian
Shuai, Bin
Zhang, Tianyi
Duan, Jingliang
Li, Shengbo Eben
Robotics
Artificial Intelligence
Machine Learning
Reinforcement learning (RL), known for its self-evolution capability, offers a promising approach to training high-level autonomous driving systems. However, handling constraints remains a significant challenge for existing RL algorithms, particularly in real-world applications. In this paper, we propose a new safety-oriented training technique called harmonic policy iteration (HPI). At each RL iteration, it first calculates two policy gradients associated with efficient driving and safety constraints, respectively. Then, a harmonic gradient is derived for policy updating, minimizing conflicts between the two gradients and consequently enabling a more balanced and stable training process. Furthermore, we adopt the state-of-the-art DSAC algorithm as the backbone and integrate it with our HPI to develop a new safe RL algorithm, DSAC-H. Extensive simulations in multi-lane scenarios demonstrate that DSAC-H achieves efficient driving performance with near-zero safety constraint violations.
title Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-lane Scenarios
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.13532