D2C-HRHR: Discrete Actions with Double Distributional Critics for High-Risk-High-Return Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jundong, Situ, Yuhui, Zhang, Fanji, Deng, Rongji, Wei, Tianqi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Soft Actor-Critic in High-Dimensional Action Spaces: The Cost of Ignoring Distribution Shift
by: Chen, Yanjun, et al.
Published: (2024)
by: Chen, Yanjun, et al.
Published: (2024)
Generative Risk Minimization for Out-of-Distribution Generalization on Graphs
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
Beyond Discrete Categories: Multi-Task Valence-Arousal Modeling for Pet Vocalization Analysis
by: Huang, Junyao, et al.
Published: (2025)
by: Huang, Junyao, et al.
Published: (2025)
Discretizing Continuous Action Space with Unimodal Probability Distributions for On-Policy Reinforcement Learning
by: Zhu, Yuanyang, et al.
Published: (2024)
by: Zhu, Yuanyang, et al.
Published: (2024)
Revisiting Discrete Soft Actor-Critic
by: Zhou, Haibin, et al.
Published: (2022)
by: Zhou, Haibin, et al.
Published: (2022)
Creative Agents: Empowering Agents with Imagination for Creative Tasks
by: Cai, Penglin, et al.
Published: (2023)
by: Cai, Penglin, et al.
Published: (2023)
SACn: Soft Actor-Critic with n-step Returns
by: Łyskawa, Jakub, et al.
Published: (2025)
by: Łyskawa, Jakub, et al.
Published: (2025)
Action-Conditioned Risk Gating for Safety-Critical Control under Partial Observability
by: Liu, Yushen, et al.
Published: (2026)
by: Liu, Yushen, et al.
Published: (2026)
Imitation Learning as Return Distribution Matching
by: Lazzati, Filippo, et al.
Published: (2025)
by: Lazzati, Filippo, et al.
Published: (2025)
Proximal Action Replacement for Behavior Cloning Actor-Critic in Offline Reinforcement Learning
by: Dong, Jinzong, et al.
Published: (2026)
by: Dong, Jinzong, et al.
Published: (2026)
Enhancing Distribution and Label Consistency for Graph Out-of-Distribution Generalization
by: Wang, Song, et al.
Published: (2025)
by: Wang, Song, et al.
Published: (2025)
HD-PiSSA: High-Rank Distributed Orthogonal Adaptation
by: Wang, Yiding, et al.
Published: (2025)
by: Wang, Yiding, et al.
Published: (2025)
Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning
by: Yang, Yuxiao, et al.
Published: (2026)
by: Yang, Yuxiao, et al.
Published: (2026)
DSAC: Distributional Soft Actor-Critic for Risk-Sensitive Reinforcement Learning
by: Ma, Xiaoteng, et al.
Published: (2020)
by: Ma, Xiaoteng, et al.
Published: (2020)
Optimizing Return Distributions with Distributional Dynamic Programming
by: Pires, Bernardo Ávila, et al.
Published: (2025)
by: Pires, Bernardo Ávila, et al.
Published: (2025)
Training-Free ANN-to-SNN Conversion for High-Performance Spiking Transformer
by: Wang, Jingya, et al.
Published: (2025)
by: Wang, Jingya, et al.
Published: (2025)
Moments Matter:Stabilizing Policy Optimization using Return Distributions
by: Jabs, Dennis, et al.
Published: (2026)
by: Jabs, Dennis, et al.
Published: (2026)
When 2D Tasks Meet 1D Serialization: On Serialization Friction in Structured Tasks
by: Lo, Chung-Hsiang, et al.
Published: (2026)
by: Lo, Chung-Hsiang, et al.
Published: (2026)
Towards Assessing and Benchmarking Risk-Return Tradeoff of Off-Policy Evaluation
by: Kiyohara, Haruka, et al.
Published: (2023)
by: Kiyohara, Haruka, et al.
Published: (2023)
Distributed Risk-Sensitive Safety Filters for Uncertain Discrete-Time Systems
by: Lederer, Armin, et al.
Published: (2025)
by: Lederer, Armin, et al.
Published: (2025)
Spectral Clustering for Discrete Distributions
by: Wang, Zixiao, et al.
Published: (2024)
by: Wang, Zixiao, et al.
Published: (2024)
GraphTOP: Graph Topology-Oriented Prompting for Graph Neural Networks
by: Fu, Xingbo, et al.
Published: (2025)
by: Fu, Xingbo, et al.
Published: (2025)
Edge Prompt Tuning for Graph Neural Networks
by: Fu, Xingbo, et al.
Published: (2025)
by: Fu, Xingbo, et al.
Published: (2025)
Generalized Discrete Diffusion with Self-Correction
by: Wang, Linxuan, et al.
Published: (2026)
by: Wang, Linxuan, et al.
Published: (2026)
Generalized Interpolating Discrete Diffusion
by: von Rütte, Dimitri, et al.
Published: (2025)
by: von Rütte, Dimitri, et al.
Published: (2025)
SAINT: Attention-Based Policies for Discrete Combinatorial Action Spaces
by: Landers, Matthew, et al.
Published: (2025)
by: Landers, Matthew, et al.
Published: (2025)
Dynamic Neighborhood Construction for Structured Large Discrete Action Spaces
by: Akkerman, Fabian, et al.
Published: (2023)
by: Akkerman, Fabian, et al.
Published: (2023)
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
by: Asad, Reza, et al.
Published: (2025)
by: Asad, Reza, et al.
Published: (2025)
Distributions as Actions: A Unified Framework for Diverse Action Spaces
by: He, Jiamin, et al.
Published: (2025)
by: He, Jiamin, et al.
Published: (2025)
In-Dataset Trajectory Return Regularization for Offline Preference-based Reinforcement Learning
by: Tu, Songjun, et al.
Published: (2024)
by: Tu, Songjun, et al.
Published: (2024)
Double Actor-Critic with TD Error-Driven Regularization in Reinforcement Learning
by: Chen, Haohui, et al.
Published: (2024)
by: Chen, Haohui, et al.
Published: (2024)
Task-Distributionally Robust Data-Free Meta-Learning
by: Hu, Zixuan, et al.
Published: (2023)
by: Hu, Zixuan, et al.
Published: (2023)
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
IDOL: Meeting Diverse Distribution Shifts with Prior Physics for Tropical Cyclone Multi-Task Estimation
by: Yan, Hanting, et al.
Published: (2025)
by: Yan, Hanting, et al.
Published: (2025)
Quality-Diversity Actor-Critic: Learning High-Performing and Diverse Behaviors via Value and Successor Features Critics
by: Grillotti, Luca, et al.
Published: (2024)
by: Grillotti, Luca, et al.
Published: (2024)
Breaking the Grid: Distance-Guided Reinforcement Learning in Large Discrete Action Spaces
by: Hoppe, Heiko, et al.
Published: (2026)
by: Hoppe, Heiko, et al.
Published: (2026)
STAS: Spatial-Temporal Return Decomposition for Multi-agent Reinforcement Learning
by: Chen, Sirui, et al.
Published: (2023)
by: Chen, Sirui, et al.
Published: (2023)
FedHERO: A Federated Learning Approach for Node Classification Task on Heterophilic Graphs
by: Chen, Zihan, et al.
Published: (2025)
by: Chen, Zihan, et al.
Published: (2025)
Beyond Affinity: A Benchmark of 1D, 2D, and 3D Methods Reveals Critical Trade-offs in Structure-Based Drug Design
by: Zheng, Kangyu, et al.
Published: (2026)
by: Zheng, Kangyu, et al.
Published: (2026)
Video Action Differencing
by: Burgess, James, et al.
Published: (2025)
by: Burgess, James, et al.
Published: (2025)
Similar Items
-
Rethinking Soft Actor-Critic in High-Dimensional Action Spaces: The Cost of Ignoring Distribution Shift
by: Chen, Yanjun, et al.
Published: (2024) -
Generative Risk Minimization for Out-of-Distribution Generalization on Graphs
by: Wang, Song, et al.
Published: (2025) -
Beyond Discrete Categories: Multi-Task Valence-Arousal Modeling for Pet Vocalization Analysis
by: Huang, Junyao, et al.
Published: (2025) -
Discretizing Continuous Action Space with Unimodal Probability Distributions for On-Policy Reinforcement Learning
by: Zhu, Yuanyang, et al.
Published: (2024) -
Revisiting Discrete Soft Actor-Critic
by: Zhou, Haibin, et al.
Published: (2022)