Moderate Actor-Critic Methods: Controlling Overestimation Bias via Expectile Loss
Fuente:
arXiv
Saved in:
| Main Authors: | Hwang, Ukjo, Hong, Songnam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Federated Reinforcement Learning in Heterogeneous Environments
by: Hwang, Ukjo, et al.
Published: (2025)
by: Hwang, Ukjo, et al.
Published: (2025)
Actor-Critic Algorithm for Dynamic Expectile and CVaR
by: Luo, Yudong, et al.
Published: (2026)
by: Luo, Yudong, et al.
Published: (2026)
Stochastic Actor-Critic: Mitigating Overestimation via Temporal Aleatoric Uncertainty
by: Özalp, Uğurcan
Published: (2026)
by: Özalp, Uğurcan
Published: (2026)
Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2024)
by: Nauman, Michal, et al.
Published: (2024)
Actor-Critic without Actor
by: Ki, Donghyeon, et al.
Published: (2025)
by: Ki, Donghyeon, et al.
Published: (2025)
Bootstrapping Expectiles in Reinforcement Learning
by: Clavier, Pierre, et al.
Published: (2024)
by: Clavier, Pierre, et al.
Published: (2024)
D2 Actor Critic: Diffusion Actor Meets Distributional Critic
by: Zhang, Lunjun, et al.
Published: (2025)
by: Zhang, Lunjun, et al.
Published: (2025)
Actor-Critic or Critic-Actor? A Tale of Two Time Scales
by: Bhatnagar, Shalabh, et al.
Published: (2022)
by: Bhatnagar, Shalabh, et al.
Published: (2022)
Actor-Critic Reinforcement Learning with Phased Actor
by: Wu, Ruofan, et al.
Published: (2024)
by: Wu, Ruofan, et al.
Published: (2024)
Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic
by: Lee, Jeong Woon, et al.
Published: (2026)
by: Lee, Jeong Woon, et al.
Published: (2026)
Human-Readable Programs as Actors of Reinforcement Learning Agents Using Critic-Moderated Evolution
by: Deproost, Senne, et al.
Published: (2024)
by: Deproost, Senne, et al.
Published: (2024)
Generative Actor Critic
by: Qin, Aoyang, et al.
Published: (2025)
by: Qin, Aoyang, et al.
Published: (2025)
Multi-Agent Soft Actor-Critic with Coordinated Loss for Autonomous Mobility-on-Demand Fleet Control
by: Woywood, Zeno, et al.
Published: (2024)
by: Woywood, Zeno, et al.
Published: (2024)
Coordinating Planning and Tracking in Layered Control Policies via Actor-Critic Learning
by: Yang, Fengjun, et al.
Published: (2024)
by: Yang, Fengjun, et al.
Published: (2024)
Hi-SAFE: Hierarchical Secure Aggregation for Lightweight Federated Learning
by: Joo, Hyeong-Gun, et al.
Published: (2025)
by: Joo, Hyeong-Gun, et al.
Published: (2025)
Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
by: Manivannan, Sanjeev, et al.
Published: (2026)
by: Manivannan, Sanjeev, et al.
Published: (2026)
A Pontryagin Method of Model-based Reinforcement Learning via Hamiltonian Actor-Critic
by: Gu, Chengyang, et al.
Published: (2026)
by: Gu, Chengyang, et al.
Published: (2026)
More Test-Time Compute Can Hurt: Overestimation Bias in LLM Beam Search
by: Dalal, Gal, et al.
Published: (2026)
by: Dalal, Gal, et al.
Published: (2026)
Neural Networks for Censored Expectile Regression Based on Data Augmentation
by: Cao, Wei, et al.
Published: (2025)
by: Cao, Wei, et al.
Published: (2025)
Mirror Descent Actor Critic via Bounded Advantage Learning
by: Iwaki, Ryo
Published: (2025)
by: Iwaki, Ryo
Published: (2025)
Finite-Time Analysis of Three-Timescale Constrained Actor-Critic and Constrained Natural Actor-Critic Algorithms
by: Panda, Prashansa, et al.
Published: (2023)
by: Panda, Prashansa, et al.
Published: (2023)
Actor-Critic Physics-informed Neural Lyapunov Control
by: Wang, Jiarui, et al.
Published: (2024)
by: Wang, Jiarui, et al.
Published: (2024)
Adaptive Ensemble Aggregation for Actor-Critics
by: Werge, Nicklas, et al.
Published: (2025)
by: Werge, Nicklas, et al.
Published: (2025)
Risk-Sensitive Exponential Actor Critic
by: Granados, Alonso, et al.
Published: (2026)
by: Granados, Alonso, et al.
Published: (2026)
Safe Langevin Soft Actor Critic
by: Keswani, Mahesh, et al.
Published: (2026)
by: Keswani, Mahesh, et al.
Published: (2026)
On the Convergence of Single-Timescale Actor-Critic
by: Kumar, Navdeep, et al.
Published: (2024)
by: Kumar, Navdeep, et al.
Published: (2024)
Actor-Critic with Active Importance Sampling
by: Molaei, Majid, et al.
Published: (2026)
by: Molaei, Majid, et al.
Published: (2026)
Broad Critic Deep Actor Reinforcement Learning for Continuous Control
by: Thalagala, Shiron, et al.
Published: (2024)
by: Thalagala, Shiron, et al.
Published: (2024)
On the Reduction of Variance and Overestimation of Deep Q-Learning
by: Sabry, Mohammed, et al.
Published: (2019)
by: Sabry, Mohammed, et al.
Published: (2019)
Deep Actor-Critics with Tight Risk Certificates
by: Tasdighi, Bahareh, et al.
Published: (2025)
by: Tasdighi, Bahareh, et al.
Published: (2025)
Efficient Soft Actor-Critic with LLM-Based Action-Level Guidance for Continuous Control
by: Ma, Hao, et al.
Published: (2026)
by: Ma, Hao, et al.
Published: (2026)
Compatible Gradient Approximations for Actor-Critic Algorithms
by: Saglam, Baturay, et al.
Published: (2024)
by: Saglam, Baturay, et al.
Published: (2024)
Refined Analysis of Entropy-Regularized Actor-Critic
by: Labbi, Safwan, et al.
Published: (2026)
by: Labbi, Safwan, et al.
Published: (2026)
Actor-Critic Pretraining for Proximal Policy Optimization
by: Kernbach, Andreas, et al.
Published: (2026)
by: Kernbach, Andreas, et al.
Published: (2026)
PAC-Bayesian Soft Actor-Critic Learning
by: Tasdighi, Bahareh, et al.
Published: (2023)
by: Tasdighi, Bahareh, et al.
Published: (2023)
Regret Analysis of Average-Reward Unichain MDPs via an Actor-Critic Approach
by: Ganesh, Swetha, et al.
Published: (2025)
by: Ganesh, Swetha, et al.
Published: (2025)
Beyond Imbalance Ratio: Data Characteristics as Critical Moderators of Oversampling Method Selection
by: Jiang, Yuwen, et al.
Published: (2026)
by: Jiang, Yuwen, et al.
Published: (2026)
Adversarial Robustness Overestimation and Instability in TRADES
by: Li, Jonathan Weiping, et al.
Published: (2024)
by: Li, Jonathan Weiping, et al.
Published: (2024)
Average-Reward Soft Actor-Critic
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
Wasserstein Barycenter Soft Actor-Critic
by: Shahrooei, Zahra, et al.
Published: (2025)
by: Shahrooei, Zahra, et al.
Published: (2025)
Similar Items
-
Federated Reinforcement Learning in Heterogeneous Environments
by: Hwang, Ukjo, et al.
Published: (2025) -
Actor-Critic Algorithm for Dynamic Expectile and CVaR
by: Luo, Yudong, et al.
Published: (2026) -
Stochastic Actor-Critic: Mitigating Overestimation via Temporal Aleatoric Uncertainty
by: Özalp, Uğurcan
Published: (2026) -
Overestimation, Overfitting, and Plasticity in Actor-Critic: the Bitter Lesson of Reinforcement Learning
by: Nauman, Michal, et al.
Published: (2024) -
Actor-Critic without Actor
by: Ki, Donghyeon, et al.
Published: (2025)