Refined Analysis of Entropy-Regularized Actor-Critic
Fuente:
arXiv
Saved in:
| Main Authors: | Labbi, Safwan, Mangold, Paul, Tiapkin, Daniil, Moulines, Eric |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled Regularization
by: Labbi, Safwan, et al.
Published: (2026)
by: Labbi, Safwan, et al.
Published: (2026)
On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments
by: Labbi, Safwan, et al.
Published: (2025)
by: Labbi, Safwan, et al.
Published: (2025)
Federated UCBVI: Communication-Efficient Federated Regret Minimization with Heterogeneous Agents
by: Labbi, Safwan, et al.
Published: (2024)
by: Labbi, Safwan, et al.
Published: (2024)
SCAFFLSA: Taming Heterogeneity in Federated Linear Stochastic Approximation and TD Learning
by: Mangold, Paul, et al.
Published: (2024)
by: Mangold, Paul, et al.
Published: (2024)
Joint Channel Selection using FedDRL in V2X
by: Mancini, Lorenzo, et al.
Published: (2024)
by: Mancini, Lorenzo, et al.
Published: (2024)
Generative Flow Networks as Entropy-Regularized RL
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
Refined Analysis of Federated Averaging and Federated Richardson-Romberg
by: Mangold, Paul, et al.
Published: (2024)
by: Mangold, Paul, et al.
Published: (2024)
Demonstration-Regularized RL
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
Improved High-Probability Bounds for the Temporal Difference Learning Algorithm via Exponential Stability
by: Samsonov, Sergey, et al.
Published: (2023)
by: Samsonov, Sergey, et al.
Published: (2023)
Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games
by: Ocello, Antonio, et al.
Published: (2025)
by: Ocello, Antonio, et al.
Published: (2025)
Convergence Guarantees for Federated SARSA with Local Training and Heterogeneous Agents
by: Mangold, Paul, et al.
Published: (2025)
by: Mangold, Paul, et al.
Published: (2025)
Scaffold with Stochastic Gradients: New Analysis with Linear Speed-Up
by: Mangold, Paul, et al.
Published: (2025)
by: Mangold, Paul, et al.
Published: (2025)
Proximal Point Nash Learning from Human Feedback
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Revisiting Mixture Policies in Entropy-Regularized Actor-Critic
by: He, Jiamin, et al.
Published: (2026)
by: He, Jiamin, et al.
Published: (2026)
Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization
by: Tiapkin, Daniil, et al.
Published: (2024)
by: Tiapkin, Daniil, et al.
Published: (2024)
Gaussian Approximation and Multiplier Bootstrap for Federated Linear Stochastic Approximation
by: Levin, Ilya, et al.
Published: (2026)
by: Levin, Ilya, et al.
Published: (2026)
Model-free Posterior Sampling via Learning Rate Randomization
by: Tiapkin, Daniil, et al.
Published: (2023)
by: Tiapkin, Daniil, et al.
Published: (2023)
ACE : Off-Policy Actor-Critic with Causality-Aware Entropy Regularization
by: Ji, Tianying, et al.
Published: (2024)
by: Ji, Tianying, et al.
Published: (2024)
Optimizing Backward Policies in GFlowNets via Trajectory Likelihood Maximization
by: Gritsaev, Timofei, et al.
Published: (2024)
by: Gritsaev, Timofei, et al.
Published: (2024)
Revisiting Non-Acyclic GFlowNets in Discrete Environments
by: Morozov, Nikita, et al.
Published: (2025)
by: Morozov, Nikita, et al.
Published: (2025)
Diffusion Actor-Critic with Entropy Regulator
by: Wang, Yinuo, et al.
Published: (2024)
by: Wang, Yinuo, et al.
Published: (2024)
Distributional Soft Actor-Critic with Three Refinements
by: Duan, Jingliang, et al.
Published: (2023)
by: Duan, Jingliang, et al.
Published: (2023)
Incentivized Learning in Principal-Agent Bandit Games
by: Scheid, Antoine, et al.
Published: (2024)
by: Scheid, Antoine, et al.
Published: (2024)
Learning Shortest Paths with Generative Flow Networks
by: Morozov, Nikita, et al.
Published: (2026)
by: Morozov, Nikita, et al.
Published: (2026)
Adaptive Set-Mass Calibration with Conformal Prediction
by: Kazantsev, Daniil, et al.
Published: (2025)
by: Kazantsev, Daniil, et al.
Published: (2025)
Maximum Entropy On-Policy Actor-Critic via Entropy Advantage Estimation
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
by: Choe, Jean Seong Bjorn, et al.
Published: (2024)
Actor-Critic without Actor
by: Ki, Donghyeon, et al.
Published: (2025)
by: Ki, Donghyeon, et al.
Published: (2025)
Improving GFlowNets with Monte Carlo Tree Search
by: Morozov, Nikita, et al.
Published: (2024)
by: Morozov, Nikita, et al.
Published: (2024)
Finite-Time Analysis of Three-Timescale Constrained Actor-Critic and Constrained Natural Actor-Critic Algorithms
by: Panda, Prashansa, et al.
Published: (2023)
by: Panda, Prashansa, et al.
Published: (2023)
DSAC-C: Constrained Maximum Entropy for Robust Discrete Soft-Actor Critic
by: Neo, Dexter, et al.
Published: (2023)
by: Neo, Dexter, et al.
Published: (2023)
Actor-Critic or Critic-Actor? A Tale of Two Time Scales
by: Bhatnagar, Shalabh, et al.
Published: (2022)
by: Bhatnagar, Shalabh, et al.
Published: (2022)
D2 Actor Critic: Diffusion Actor Meets Distributional Critic
by: Zhang, Lunjun, et al.
Published: (2025)
by: Zhang, Lunjun, et al.
Published: (2025)
Actor-Critic Reinforcement Learning with Phased Actor
by: Wu, Ruofan, et al.
Published: (2024)
by: Wu, Ruofan, et al.
Published: (2024)
Tight Analysis of Decentralized SGD: A Markov Chain Perspective
by: Versini, Lucas, et al.
Published: (2026)
by: Versini, Lucas, et al.
Published: (2026)
Double Actor-Critic with TD Error-Driven Regularization in Reinforcement Learning
by: Chen, Haohui, et al.
Published: (2024)
by: Chen, Haohui, et al.
Published: (2024)
Generative Actor Critic
by: Qin, Aoyang, et al.
Published: (2025)
by: Qin, Aoyang, et al.
Published: (2025)
ARAC: Adaptive Regularized Multi-Agent Soft Actor-Critic in Graph-Structured Adversarial Games
by: Shi, Ruochuan, et al.
Published: (2025)
by: Shi, Ruochuan, et al.
Published: (2025)
Balanced Training of Energy-Based Models with Adaptive Flow Sampling
by: Grenioux, Louis, et al.
Published: (2023)
by: Grenioux, Louis, et al.
Published: (2023)
gfnx: Fast and Scalable Library for Generative Flow Networks in JAX
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Risk-Sensitive Exponential Actor Critic
by: Granados, Alonso, et al.
Published: (2026)
by: Granados, Alonso, et al.
Published: (2026)
Similar Items
-
Beyond Softmax and Entropy: Convergence Rates of Policy Gradients with f-SoftArgmax Parameterization & Coupled Regularization
by: Labbi, Safwan, et al.
Published: (2026) -
On Global Convergence Rates for Federated Softmax Policy Gradient under Heterogeneous Environments
by: Labbi, Safwan, et al.
Published: (2025) -
Federated UCBVI: Communication-Efficient Federated Regret Minimization with Heterogeneous Agents
by: Labbi, Safwan, et al.
Published: (2024) -
SCAFFLSA: Taming Heterogeneity in Federated Linear Stochastic Approximation and TD Learning
by: Mangold, Paul, et al.
Published: (2024) -
Joint Channel Selection using FedDRL in V2X
by: Mancini, Lorenzo, et al.
Published: (2024)