D2C-HRHR: Discrete Actions with Double Distributional Critics for High-Risk-High-Return Tasks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Jundong, Situ, Yuhui, Zhang, Fanji, Deng, Rongji, Wei, Tianqi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918164018757632
author Zhang, Jundong
Situ, Yuhui
Zhang, Fanji
Deng, Rongji
Wei, Tianqi
author_facet Zhang, Jundong
Situ, Yuhui
Zhang, Fanji
Deng, Rongji
Wei, Tianqi
contents Tasks involving high-risk-high-return (HRHR) actions, such as obstacle crossing, often exhibit multimodal action distributions and stochastic returns. Most reinforcement learning (RL) methods assume unimodal Gaussian policies and rely on scalar-valued critics, which limits their effectiveness in HRHR settings. We formally define HRHR tasks and theoretically show that Gaussian policies cannot guarantee convergence to the optimal solution. To address this, we propose a reinforcement learning framework that (i) discretizes continuous action spaces to approximate multimodal distributions, (ii) employs entropy-regularized exploration to improve coverage of risky but rewarding actions, and (iii) introduces a dual-critic architecture for more accurate discrete value distribution estimation. The framework scales to high-dimensional action spaces, supporting complex control domains. Experiments on locomotion and manipulation benchmarks with high risks of failure demonstrate that our method outperforms baselines, underscoring the importance of explicitly modeling multimodality and risk in RL.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17212
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle D2C-HRHR: Discrete Actions with Double Distributional Critics for High-Risk-High-Return Tasks
Zhang, Jundong
Situ, Yuhui
Zhang, Fanji
Deng, Rongji
Wei, Tianqi
Machine Learning
Artificial Intelligence
Tasks involving high-risk-high-return (HRHR) actions, such as obstacle crossing, often exhibit multimodal action distributions and stochastic returns. Most reinforcement learning (RL) methods assume unimodal Gaussian policies and rely on scalar-valued critics, which limits their effectiveness in HRHR settings. We formally define HRHR tasks and theoretically show that Gaussian policies cannot guarantee convergence to the optimal solution. To address this, we propose a reinforcement learning framework that (i) discretizes continuous action spaces to approximate multimodal distributions, (ii) employs entropy-regularized exploration to improve coverage of risky but rewarding actions, and (iii) introduces a dual-critic architecture for more accurate discrete value distribution estimation. The framework scales to high-dimensional action spaces, supporting complex control domains. Experiments on locomotion and manipulation benchmarks with high risks of failure demonstrate that our method outperforms baselines, underscoring the importance of explicitly modeling multimodality and risk in RL.
title D2C-HRHR: Discrete Actions with Double Distributional Critics for High-Risk-High-Return Tasks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.17212