Revisiting Mixture Policies in Entropy-Regularized Actor-Critic
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Jiamin, Neumann, Samuel, Mei, Jincheng, White, Adam, White, Martha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Empirical Design in Reinforcement Learning
von: Patterson, Andrew, et al.
Veröffentlicht: (2023)
von: Patterson, Andrew, et al.
Veröffentlicht: (2023)
ACE : Off-Policy Actor-Critic with Causality-Aware Entropy Regularization
von: Ji, Tianying, et al.
Veröffentlicht: (2024)
von: Ji, Tianying, et al.
Veröffentlicht: (2024)
Distributions as Actions: A Unified Framework for Diverse Action Spaces
von: He, Jiamin, et al.
Veröffentlicht: (2025)
von: He, Jiamin, et al.
Veröffentlicht: (2025)
Maximum Entropy On-Policy Actor-Critic via Entropy Advantage Estimation
von: Choe, Jean Seong Bjorn, et al.
Veröffentlicht: (2024)
von: Choe, Jean Seong Bjorn, et al.
Veröffentlicht: (2024)
Symmetric Behavior Regularized Policy Optimization
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
Fine-Tuning without Performance Degradation
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
von: Patterson, Andrew, et al.
Veröffentlicht: (2021)
von: Patterson, Andrew, et al.
Veröffentlicht: (2021)
Revisiting Discrete Soft Actor-Critic
von: Zhou, Haibin, et al.
Veröffentlicht: (2022)
von: Zhou, Haibin, et al.
Veröffentlicht: (2022)
Diffusion Actor-Critic with Entropy Regulator
von: Wang, Yinuo, et al.
Veröffentlicht: (2024)
von: Wang, Yinuo, et al.
Veröffentlicht: (2024)
Real-Time Recurrent Learning using Trace Units in Reinforcement Learning
von: Elelimy, Esraa, et al.
Veröffentlicht: (2024)
von: Elelimy, Esraa, et al.
Veröffentlicht: (2024)
Investigating the Interplay of Prioritized Replay and Generalization
von: Panahi, Parham Mohammad, et al.
Veröffentlicht: (2024)
von: Panahi, Parham Mohammad, et al.
Veröffentlicht: (2024)
Forager: a lightweight testbed for continual learning with partial observability in RL
von: Tang, Steven, et al.
Veröffentlicht: (2026)
von: Tang, Steven, et al.
Veröffentlicht: (2026)
Distributional Soft Actor-Critic with Diffusion Policy
von: Liu, Tong, et al.
Veröffentlicht: (2025)
von: Liu, Tong, et al.
Veröffentlicht: (2025)
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
von: Liu, Vincent, et al.
Veröffentlicht: (2023)
von: Liu, Vincent, et al.
Veröffentlicht: (2023)
A New View on Planning in Online Reinforcement Learning
von: Roice, Kevin, et al.
Veröffentlicht: (2024)
von: Roice, Kevin, et al.
Veröffentlicht: (2024)
Double Actor-Critic with TD Error-Driven Regularization in Reinforcement Learning
von: Chen, Haohui, et al.
Veröffentlicht: (2024)
von: Chen, Haohui, et al.
Veröffentlicht: (2024)
Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic
von: Lee, Jeong Woon, et al.
Veröffentlicht: (2026)
von: Lee, Jeong Woon, et al.
Veröffentlicht: (2026)
Deep Reinforcement Learning with Gradient Eligibility Traces
von: Elelimy, Esraa, et al.
Veröffentlicht: (2025)
von: Elelimy, Esraa, et al.
Veröffentlicht: (2025)
Functional Critics Are Essential for Actor-Critic: From Off-Policy Stability to Efficient Exploration
von: Bai, Qinxun, et al.
Veröffentlicht: (2025)
von: Bai, Qinxun, et al.
Veröffentlicht: (2025)
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
von: Vasan, Gautham, et al.
Veröffentlicht: (2024)
von: Vasan, Gautham, et al.
Veröffentlicht: (2024)
Limits of Actor-Critic Algorithms for Decision Tree Policies Learning in IBMDPs
von: Kohler, Hector, et al.
Veröffentlicht: (2023)
von: Kohler, Hector, et al.
Veröffentlicht: (2023)
Enabling Off-Policy Imitation Learning with Deep Actor Critic Stabilization
von: Sen, Sayambhu, et al.
Veröffentlicht: (2025)
von: Sen, Sayambhu, et al.
Veröffentlicht: (2025)
Studying the Interplay Between the Actor and Critic Representations in Reinforcement Learning
von: Garcin, Samuel, et al.
Veröffentlicht: (2025)
von: Garcin, Samuel, et al.
Veröffentlicht: (2025)
SMAC: Score-Matched Actor-Critics for Robust Offline-to-Online Transfer
von: de Lara, Nathan Samuel, et al.
Veröffentlicht: (2026)
von: de Lara, Nathan Samuel, et al.
Veröffentlicht: (2026)
Collaborative Yet Personalized Policy Training: Single-Timescale Federated Actor-Critic
von: Wang, Leo Muxing, et al.
Veröffentlicht: (2026)
von: Wang, Leo Muxing, et al.
Veröffentlicht: (2026)
Causal Policy Learning in Reinforcement Learning: Backdoor-Adjusted Soft Actor-Critic
von: Vo, Thanh Vinh, et al.
Veröffentlicht: (2025)
von: Vo, Thanh Vinh, et al.
Veröffentlicht: (2025)
Adaptive Horizon Actor-Critic for Policy Learning in Contact-Rich Differentiable Simulation
von: Georgiev, Ignat, et al.
Veröffentlicht: (2024)
von: Georgiev, Ignat, et al.
Veröffentlicht: (2024)
Relative Importance Sampling for off-Policy Actor-Critic in Deep Reinforcement Learning
von: Humayoo, Mahammad, et al.
Veröffentlicht: (2018)
von: Humayoo, Mahammad, et al.
Veröffentlicht: (2018)
Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2025)
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2025)
Value Bonuses using Ensemble Errors for Exploration in Reinforcement Learning
von: Wahab, Abdul, et al.
Veröffentlicht: (2026)
von: Wahab, Abdul, et al.
Veröffentlicht: (2026)
Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models
von: Jajoo, Pranaya, et al.
Veröffentlicht: (2026)
von: Jajoo, Pranaya, et al.
Veröffentlicht: (2026)
Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
von: Manivannan, Sanjeev, et al.
Veröffentlicht: (2026)
von: Manivannan, Sanjeev, et al.
Veröffentlicht: (2026)
Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning
von: Yang, Tong, et al.
Veröffentlicht: (2023)
von: Yang, Tong, et al.
Veröffentlicht: (2023)
Gradient Iterated Temporal-Difference Learning
von: Vincent, Théo, et al.
Veröffentlicht: (2026)
von: Vincent, Théo, et al.
Veröffentlicht: (2026)
Goal-Space Planning with Subgoal Models
von: Lo, Chunlok, et al.
Veröffentlicht: (2022)
von: Lo, Chunlok, et al.
Veröffentlicht: (2022)
Value Improved Actor Critic Algorithms
von: Oren, Yaniv, et al.
Veröffentlicht: (2024)
von: Oren, Yaniv, et al.
Veröffentlicht: (2024)
Average-Reward Soft Actor-Critic
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
Deep Double Q-learning
von: Nagarajan, Prabhat, et al.
Veröffentlicht: (2025)
von: Nagarajan, Prabhat, et al.
Veröffentlicht: (2025)
Demystifying the Recency Heuristic in Temporal-Difference Learning
von: Daley, Brett, et al.
Veröffentlicht: (2024)
von: Daley, Brett, et al.
Veröffentlicht: (2024)
Variational Distillation of Diffusion Policies into Mixture of Experts
von: Zhou, Hongyi, et al.
Veröffentlicht: (2024)
von: Zhou, Hongyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Empirical Design in Reinforcement Learning
von: Patterson, Andrew, et al.
Veröffentlicht: (2023) -
ACE : Off-Policy Actor-Critic with Causality-Aware Entropy Regularization
von: Ji, Tianying, et al.
Veröffentlicht: (2024) -
Distributions as Actions: A Unified Framework for Diverse Action Spaces
von: He, Jiamin, et al.
Veröffentlicht: (2025) -
Maximum Entropy On-Policy Actor-Critic via Entropy Advantage Estimation
von: Choe, Jean Seong Bjorn, et al.
Veröffentlicht: (2024) -
Symmetric Behavior Regularized Policy Optimization
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)