Two-Timescale Critic-Actor for Average Reward MDPs with Function Approximation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Panda, Prashansa, Bhatnagar, Shalabh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Finite-Time Analysis of Three-Timescale Constrained Actor-Critic and Constrained Natural Actor-Critic Algorithms
von: Panda, Prashansa, et al.
Veröffentlicht: (2023)
von: Panda, Prashansa, et al.
Veröffentlicht: (2023)
Finite Time Analysis of Constrained Natural Critic-Actor Algorithm with Improved Sample Complexity
von: Panda, Prashansa, et al.
Veröffentlicht: (2025)
von: Panda, Prashansa, et al.
Veröffentlicht: (2025)
Actor-Critic or Critic-Actor? A Tale of Two Time Scales
von: Bhatnagar, Shalabh, et al.
Veröffentlicht: (2022)
von: Bhatnagar, Shalabh, et al.
Veröffentlicht: (2022)
An Actor-Critic Algorithm with Function Approximation for Risk Sensitive Cost Markov Decision Processes
von: Guin, Soumyajit, et al.
Veröffentlicht: (2025)
von: Guin, Soumyajit, et al.
Veröffentlicht: (2025)
Enabling Off-Policy Imitation Learning with Deep Actor Critic Stabilization
von: Sen, Sayambhu, et al.
Veröffentlicht: (2025)
von: Sen, Sayambhu, et al.
Veröffentlicht: (2025)
Regret Analysis of Average-Reward Unichain MDPs via an Actor-Critic Approach
von: Ganesh, Swetha, et al.
Veröffentlicht: (2025)
von: Ganesh, Swetha, et al.
Veröffentlicht: (2025)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
von: Xu, Yang, et al.
Veröffentlicht: (2025)
von: Xu, Yang, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
Average-Reward Soft Actor-Critic
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
Primal-Only Actor Critic Algorithm for Robust Constrained Average Cost MDPs
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2025)
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2025)
On the Convergence of Single-Timescale Actor-Critic
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
Deep SOR Minimax Q-learning for Two-player Zero-sum Game
von: Gautam, Saksham, et al.
Veröffentlicht: (2025)
von: Gautam, Saksham, et al.
Veröffentlicht: (2025)
A policy gradient approach for Finite Horizon Constrained Markov Decision Processes
von: Guin, Soumyajit, et al.
Veröffentlicht: (2022)
von: Guin, Soumyajit, et al.
Veröffentlicht: (2022)
Convergent Reinforcement Learning Algorithms for Stochastic Shortest Path Problem
von: Guin, Soumyajit, et al.
Veröffentlicht: (2025)
von: Guin, Soumyajit, et al.
Veröffentlicht: (2025)
Finite-time analysis of Multi-timescale Stochastic Optimization Algorithms
von: Kartikey, Kaustubh, et al.
Veröffentlicht: (2026)
von: Kartikey, Kaustubh, et al.
Veröffentlicht: (2026)
n-Step Temporal Difference Learning with Optimal n
von: Mandal, Lakshmi, et al.
Veröffentlicht: (2023)
von: Mandal, Lakshmi, et al.
Veröffentlicht: (2023)
Approximate Linear Programming for Decentralized Policy Iteration in Cooperative Multi-agent Markov Decision Processes
von: Mandal, Lakshmi, et al.
Veröffentlicht: (2023)
von: Mandal, Lakshmi, et al.
Veröffentlicht: (2023)
Global Convergence of Average Reward Constrained MDPs with Neural Critic and General Policy Parameterization
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2026)
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2026)
Multi Timescale Stochastic Approximation: Stability and Convergence
von: Deb, Rohan, et al.
Veröffentlicht: (2021)
von: Deb, Rohan, et al.
Veröffentlicht: (2021)
Convergence of Multiagent Learning Systems for Traffic control
von: Sen, Sayambhu, et al.
Veröffentlicht: (2025)
von: Sen, Sayambhu, et al.
Veröffentlicht: (2025)
Sample-efficient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation
von: He, Jianliang, et al.
Veröffentlicht: (2024)
von: He, Jianliang, et al.
Veröffentlicht: (2024)
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
von: Wei, Yukuan, et al.
Veröffentlicht: (2025)
von: Wei, Yukuan, et al.
Veröffentlicht: (2025)
Regret Analysis of Unichain Average Reward Constrained MDPs with General Parameterization
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2026)
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2026)
Efficient $Q$-Learning and Actor-Critic Methods for Robust Average Reward Reinforcement Learning
von: Xu, Yang, et al.
Veröffentlicht: (2025)
von: Xu, Yang, et al.
Veröffentlicht: (2025)
Breaking the Computational Barrier: Provably Efficient Actor-Critic for Low-Rank MDPs
von: Huang, Ruiquan, et al.
Veröffentlicht: (2026)
von: Huang, Ruiquan, et al.
Veröffentlicht: (2026)
A Sharper Global Convergence Analysis for Average Reward Reinforcement Learning via an Actor-Critic Approach
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
von: Boone, Victor, et al.
Veröffentlicht: (2024)
von: Boone, Victor, et al.
Veröffentlicht: (2024)
Span-Based Optimal Sample Complexity for Average Reward MDPs
von: Zurek, Matthew, et al.
Veröffentlicht: (2023)
von: Zurek, Matthew, et al.
Veröffentlicht: (2023)
Collaborative Yet Personalized Policy Training: Single-Timescale Federated Actor-Critic
von: Wang, Leo Muxing, et al.
Veröffentlicht: (2026)
von: Wang, Leo Muxing, et al.
Veröffentlicht: (2026)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2025)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2025)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
von: Wang, Shengbo, et al.
Veröffentlicht: (2026)
von: Wang, Shengbo, et al.
Veröffentlicht: (2026)
Heavy-Ball Momentum Accelerated Actor-Critic With Function Approximation
von: Dong, Yanjie, et al.
Veröffentlicht: (2024)
von: Dong, Yanjie, et al.
Veröffentlicht: (2024)
Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
Compatible Gradient Approximations for Actor-Critic Algorithms
von: Saglam, Baturay, et al.
Veröffentlicht: (2024)
von: Saglam, Baturay, et al.
Veröffentlicht: (2024)
Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
von: Manivannan, Sanjeev, et al.
Veröffentlicht: (2026)
von: Manivannan, Sanjeev, et al.
Veröffentlicht: (2026)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
von: Zurek, Matthew, et al.
Veröffentlicht: (2024)
von: Zurek, Matthew, et al.
Veröffentlicht: (2024)
Generalized Random Direction Newton Algorithms for Stochastic Optimization
von: Pachal, Soumen, et al.
Veröffentlicht: (2026)
von: Pachal, Soumen, et al.
Veröffentlicht: (2026)
Probabilistic Safety Guarantee for Stochastic Control Systems Using Average Reward MDPs
von: Omidi, Saber, et al.
Veröffentlicht: (2025)
von: Omidi, Saber, et al.
Veröffentlicht: (2025)
Is Pure Exploitation Sufficient in Exogenous MDPs with Linear Function Approximation?
von: Liang, Hao, et al.
Veröffentlicht: (2026)
von: Liang, Hao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Finite-Time Analysis of Three-Timescale Constrained Actor-Critic and Constrained Natural Actor-Critic Algorithms
von: Panda, Prashansa, et al.
Veröffentlicht: (2023) -
Finite Time Analysis of Constrained Natural Critic-Actor Algorithm with Improved Sample Complexity
von: Panda, Prashansa, et al.
Veröffentlicht: (2025) -
Actor-Critic or Critic-Actor? A Tale of Two Time Scales
von: Bhatnagar, Shalabh, et al.
Veröffentlicht: (2022) -
An Actor-Critic Algorithm with Function Approximation for Risk Sensitive Cost Markov Decision Processes
von: Guin, Soumyajit, et al.
Veröffentlicht: (2025) -
Enabling Off-Policy Imitation Learning with Deep Actor Critic Stabilization
von: Sen, Sayambhu, et al.
Veröffentlicht: (2025)