Convergence of actor-critic for entropy regularised MDPs in general action spaces
Fuente:
arXiv
Saved in:
| Main Authors: | Zorba, Denis, Šiška, David, Szpruch, Lukasz |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mirror descent actor-critic methods for entropy regularised MDPs in general spaces: stability and convergence
by: Zorba, Denis, et al.
Published: (2026)
by: Zorba, Denis, et al.
Published: (2026)
A Fisher-Rao gradient flow for entropy-regularised Markov decision processes in Polish spaces
by: Kerimkulov, Bekzhan, et al.
Published: (2023)
by: Kerimkulov, Bekzhan, et al.
Published: (2023)
PPO in the Fisher-Rao geometry
by: Lascu, Razvan-Andrei, et al.
Published: (2025)
by: Lascu, Razvan-Andrei, et al.
Published: (2025)
Gradient Flows for Regularized Stochastic Control Problems
by: Šiška, David, et al.
Published: (2020)
by: Šiška, David, et al.
Published: (2020)
Linear convergence of proximal descent schemes on the Wasserstein space
by: Lascu, Razvan-Andrei, et al.
Published: (2024)
by: Lascu, Razvan-Andrei, et al.
Published: (2024)
Mirror Descent for Stochastic Control Problems with Measure-valued Controls
by: Kerimkulov, Bekzhan, et al.
Published: (2024)
by: Kerimkulov, Bekzhan, et al.
Published: (2024)
Logarithmic regret in the ergodic Avellaneda-Stoikov market making model
by: Cao, Jialun, et al.
Published: (2024)
by: Cao, Jialun, et al.
Published: (2024)
Entropic mean-field min-max problems via Best Response flow
by: Lascu, Razvan-Andrei, et al.
Published: (2023)
by: Lascu, Razvan-Andrei, et al.
Published: (2023)
Mirror descent for constrained stochastic control problems
by: Sethi, Deven, et al.
Published: (2025)
by: Sethi, Deven, et al.
Published: (2025)
Mirror Descent-Ascent for mean-field min-max problems
by: Lascu, Razvan-Andrei, et al.
Published: (2024)
by: Lascu, Razvan-Andrei, et al.
Published: (2024)
A Fisher-Rao gradient flow for entropic mean-field min-max games
by: Lascu, Razvan-Andrei, et al.
Published: (2024)
by: Lascu, Razvan-Andrei, et al.
Published: (2024)
A note on convergence of Wasserstein policy optimization
by: Šiška, David, et al.
Published: (2026)
by: Šiška, David, et al.
Published: (2026)
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
by: Ding, Dongsheng, et al.
Published: (2023)
by: Ding, Dongsheng, et al.
Published: (2023)
A factorisation-based regularised interior point method using the augmented system
by: Zanetti, Filippo, et al.
Published: (2025)
by: Zanetti, Filippo, et al.
Published: (2025)
Off-the-grid regularisation for Poisson inverse problems
by: Lazzaretti, Marta, et al.
Published: (2024)
by: Lazzaretti, Marta, et al.
Published: (2024)
Dynamic inverse problems: Online regularisation theory
by: Jauhiainen, Jyrki, et al.
Published: (2026)
by: Jauhiainen, Jyrki, et al.
Published: (2026)
$ε$-Policy Gradient for Online Pricing
by: Szpruch, Lukasz, et al.
Published: (2024)
by: Szpruch, Lukasz, et al.
Published: (2024)
Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
by: Ding, Dongsheng, et al.
Published: (2022)
by: Ding, Dongsheng, et al.
Published: (2022)
Constrained Average-Reward Intermittently Observable MDPs
by: Avrachenkov, Konstantin, et al.
Published: (2025)
by: Avrachenkov, Konstantin, et al.
Published: (2025)
An adaptive cubic regularisation algorithm based on interior point methods for solving nonlinear inequality constrained optimization
by: Pei, Yonggang, et al.
Published: (2024)
by: Pei, Yonggang, et al.
Published: (2024)
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
by: Zhang, Runyu, et al.
Published: (2023)
by: Zhang, Runyu, et al.
Published: (2023)
Epsilon-Optimal Policies for Average-Cost Separable MDPs with Perturbations
by: Kantawala, Dhairya
Published: (2025)
by: Kantawala, Dhairya
Published: (2025)
Finite-time analysis of single-timescale actor-critic
by: Chen, Xuyang, et al.
Published: (2022)
by: Chen, Xuyang, et al.
Published: (2022)
Extended-variable relaxations for the constrained generalized maximum-entropy sampling problem
by: Ponte, Gabriel, et al.
Published: (2026)
by: Ponte, Gabriel, et al.
Published: (2026)
Data-Driven Non-Parametric Model Learning and Adaptive Control of MDPs with Borel spaces: Identifiability and Near Optimal Design
by: Mrani-Zentar, Omar, et al.
Published: (2025)
by: Mrani-Zentar, Omar, et al.
Published: (2025)
Sequential Decision-Making under Uncertainty: A Robust MDPs review
by: Ou, Wenfan, et al.
Published: (2024)
by: Ou, Wenfan, et al.
Published: (2024)
Revisiting Subgradient Dominance in Robust MDPs: Counterexamples, Hardness, and Sufficient Conditions
by: Kitamura, Toshinori, et al.
Published: (2026)
by: Kitamura, Toshinori, et al.
Published: (2026)
Model approximation in MDPs with unbounded per-step cost
by: Bozkurt, Berk, et al.
Published: (2024)
by: Bozkurt, Berk, et al.
Published: (2024)
Linear Dynamics meets Linear MDPs: Closed-Form Optimal Policies via Reinforcement Learning
by: Makdah, Abed AlRahman Al, et al.
Published: (2025)
by: Makdah, Abed AlRahman Al, et al.
Published: (2025)
Robustness to Model Approximation, Model Learning From Data, and Sample Complexity in Wasserstein Regular MDPs
by: Zhou, Yichen, et al.
Published: (2024)
by: Zhou, Yichen, et al.
Published: (2024)
Faster Fixed-Point Methods for Multichain MDPs
by: Zurek, Matthew, et al.
Published: (2025)
by: Zurek, Matthew, et al.
Published: (2025)
Planning and Learning in Average Risk-aware MDPs
by: Wang, Weikai, et al.
Published: (2025)
by: Wang, Weikai, et al.
Published: (2025)
Convergence of linesearch-based generalized conditional gradient methods without smoothness assumptions
by: Yagishita, Shotaro
Published: (2025)
by: Yagishita, Shotaro
Published: (2025)
Convergence guarantees for stochastic algorithms solving non-unique problems in metric spaces
by: Pischke, Nicholas, et al.
Published: (2026)
by: Pischke, Nicholas, et al.
Published: (2026)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
by: Boone, Victor, et al.
Published: (2024)
by: Boone, Victor, et al.
Published: (2024)
Generalized specific entropy on Wiener space with application to Martingale Optimal Transport
by: Buet-Golfouse, Francois, et al.
Published: (2026)
by: Buet-Golfouse, Francois, et al.
Published: (2026)
Convex relaxation for the generalized maximum-entropy sampling problem
by: Ponte, Gabriel, et al.
Published: (2024)
by: Ponte, Gabriel, et al.
Published: (2024)
Second-order methods for quartically-regularised cubic polynomials, with applications to high-order tensor methods
by: Cartis, Coralia, et al.
Published: (2023)
by: Cartis, Coralia, et al.
Published: (2023)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
by: Zhang, Zhongjun, et al.
Published: (2026)
by: Zhang, Zhongjun, et al.
Published: (2026)
Efficient Model-Free Exploration in Low-Rank MDPs
by: Mhammedi, Zakaria, et al.
Published: (2023)
by: Mhammedi, Zakaria, et al.
Published: (2023)
Similar Items
-
Mirror descent actor-critic methods for entropy regularised MDPs in general spaces: stability and convergence
by: Zorba, Denis, et al.
Published: (2026) -
A Fisher-Rao gradient flow for entropy-regularised Markov decision processes in Polish spaces
by: Kerimkulov, Bekzhan, et al.
Published: (2023) -
PPO in the Fisher-Rao geometry
by: Lascu, Razvan-Andrei, et al.
Published: (2025) -
Gradient Flows for Regularized Stochastic Control Problems
by: Šiška, David, et al.
Published: (2020) -
Linear convergence of proximal descent schemes on the Wasserstein space
by: Lascu, Razvan-Andrei, et al.
Published: (2024)