Policy Gradient Methods for Non-Markovian Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Kar, Avik, Chandak, Siddharth, Singh, Rahul, Sinhahajari, Soumitra, Moulines, Eric, Bhatnagar, Shalabh, Bambos, Nicholas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
High-Probability Bounds for SGD under the Polyak-Lojasiewicz Condition with Markovian Noise
by: Kar, Avik, et al.
Published: (2026)
by: Kar, Avik, et al.
Published: (2026)
Regret and Sample Complexity of Online Q-Learning via Concentration of Stochastic Approximation with Time-Inhomogeneous Markov Chains
by: Singh, Rahul, et al.
Published: (2026)
by: Singh, Rahul, et al.
Published: (2026)
Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
by: Kar, Avik, et al.
Published: (2024)
by: Kar, Avik, et al.
Published: (2024)
Enabling Off-Policy Imitation Learning with Deep Actor Critic Stabilization
by: Sen, Sayambhu, et al.
Published: (2025)
by: Sen, Sayambhu, et al.
Published: (2025)
Finite-Time Bounds for Two-Time-Scale Stochastic Approximation with Arbitrary Norm Contractions and Markovian Noise
by: Chandak, Siddharth, et al.
Published: (2025)
by: Chandak, Siddharth, et al.
Published: (2025)
Convergence of Multiagent Learning Systems for Traffic control
by: Sen, Sayambhu, et al.
Published: (2025)
by: Sen, Sayambhu, et al.
Published: (2025)
Finite Time Analysis of Constrained Natural Critic-Actor Algorithm with Improved Sample Complexity
by: Panda, Prashansa, et al.
Published: (2025)
by: Panda, Prashansa, et al.
Published: (2025)
Policy Zooming: Adaptive Discretization-based Infinite-Horizon Average-Reward Reinforcement Learning
by: Kar, Avik, et al.
Published: (2024)
by: Kar, Avik, et al.
Published: (2024)
Learning to Control Unknown Strongly Monotone Games
by: Chandak, Siddharth, et al.
Published: (2024)
by: Chandak, Siddharth, et al.
Published: (2024)
Last-Iterate Guarantees for Learning in Co-coercive Games
by: Chandak, Siddharth, et al.
Published: (2026)
by: Chandak, Siddharth, et al.
Published: (2026)
Choose Your Battles: Distributed Learning Over Multiple Tug of War Games
by: Chandak, Siddharth, et al.
Published: (2025)
by: Chandak, Siddharth, et al.
Published: (2025)
Reinforcement Learning in Non-Markovian Environments
by: Chandak, Siddharth, et al.
Published: (2022)
by: Chandak, Siddharth, et al.
Published: (2022)
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
by: Barakat, Anas, et al.
Published: (2024)
by: Barakat, Anas, et al.
Published: (2024)
Heavy-Tailed and Long-Range Dependent Noise in Stochastic Approximation: A Finite-Time Analysis
by: Chandak, Siddharth, et al.
Published: (2026)
by: Chandak, Siddharth, et al.
Published: (2026)
PG-Rainbow: Using Distributional Reinforcement Learning in Policy Gradient Methods
by: Jeon, WooJae, et al.
Published: (2024)
by: Jeon, WooJae, et al.
Published: (2024)
Safe Reinforcement Learning with Learned Non-Markovian Safety Constraints
by: Low, Siow Meng, et al.
Published: (2024)
by: Low, Siow Meng, et al.
Published: (2024)
Policy Dispersion in Non-Markovian Environment
by: Qu, Bohao, et al.
Published: (2023)
by: Qu, Bohao, et al.
Published: (2023)
Convergent Reinforcement Learning Algorithms for Stochastic Shortest Path Problem
by: Guin, Soumyajit, et al.
Published: (2025)
by: Guin, Soumyajit, et al.
Published: (2025)
Learning General Policies with Policy Gradient Methods
by: Ståhlberg, Simon, et al.
Published: (2025)
by: Ståhlberg, Simon, et al.
Published: (2025)
The ODE Method for Stochastic Approximation and Reinforcement Learning with Markovian Noise
by: Liu, Shuze Daniel, et al.
Published: (2024)
by: Liu, Shuze Daniel, et al.
Published: (2024)
Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning
by: Yang, Tong, et al.
Published: (2023)
by: Yang, Tong, et al.
Published: (2023)
Spatial-Temporal Reinforcement Learning for Network Routing with Non-Markovian Traffic
by: Wang, Molly, et al.
Published: (2025)
by: Wang, Molly, et al.
Published: (2025)
Policy Gradients for Cumulative Prospect Theory in Reinforcement Learning
by: Lepel, Olivier, et al.
Published: (2024)
by: Lepel, Olivier, et al.
Published: (2024)
Model-Based Reinforcement Learning in Discrete-Action Non-Markovian Reward Decision Processes
by: Trapasso, Alessandro, et al.
Published: (2025)
by: Trapasso, Alessandro, et al.
Published: (2025)
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence
by: Xiao, Minheng, et al.
Published: (2024)
by: Xiao, Minheng, et al.
Published: (2024)
Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
by: Nishimori, Soichiro, et al.
Published: (2026)
by: Nishimori, Soichiro, et al.
Published: (2026)
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
by: Reddi, Aryaman, et al.
Published: (2025)
by: Reddi, Aryaman, et al.
Published: (2025)
Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement Learning
by: Batra, Sumeet, et al.
Published: (2023)
by: Batra, Sumeet, et al.
Published: (2023)
CAPSULE: Control-Theoretic Action Perturbations for Safe Uncertainty-Aware Reinforcement Learning
by: Narava, Rahul, et al.
Published: (2026)
by: Narava, Rahul, et al.
Published: (2026)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
by: Melo, Luckeciano C., et al.
Published: (2025)
by: Melo, Luckeciano C., et al.
Published: (2025)
VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory
by: Lei, Yuheng, et al.
Published: (2026)
by: Lei, Yuheng, et al.
Published: (2026)
Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement Learning
by: Shitanda, Naoki, et al.
Published: (2026)
by: Shitanda, Naoki, et al.
Published: (2026)
LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem Exploration
by: Qiu, Ruiyu, et al.
Published: (2025)
by: Qiu, Ruiyu, et al.
Published: (2025)
The Definitive Guide to Policy Gradients in Deep Reinforcement Learning: Theory, Algorithms and Implementations
by: Lehmann, Matthias
Published: (2024)
by: Lehmann, Matthias
Published: (2024)
Fast and Robust Likelihood-Guided Diffusion Posterior Sampling with Amortized Variational Inference
by: Zheng, Léon, et al.
Published: (2026)
by: Zheng, Léon, et al.
Published: (2026)
Logit Dynamics in Softmax Policy Gradient Methods
by: Li, Yingru
Published: (2025)
by: Li, Yingru
Published: (2025)
Performative Policy Gradient: Optimality in Performative Reinforcement Learning
by: Basu, Debabrota, et al.
Published: (2025)
by: Basu, Debabrota, et al.
Published: (2025)
n-Step Temporal Difference Learning with Optimal n
by: Mandal, Lakshmi, et al.
Published: (2023)
by: Mandal, Lakshmi, et al.
Published: (2023)
Rank-1 Approximation of Inverse Fisher for Natural Policy Gradients in Deep Reinforcement Learning
by: Huo, Yingxiao, et al.
Published: (2026)
by: Huo, Yingxiao, et al.
Published: (2026)
Mollification Effects of Policy Gradient Methods
by: Wang, Tao, et al.
Published: (2024)
by: Wang, Tao, et al.
Published: (2024)
Similar Items
-
High-Probability Bounds for SGD under the Polyak-Lojasiewicz Condition with Markovian Noise
by: Kar, Avik, et al.
Published: (2026) -
Regret and Sample Complexity of Online Q-Learning via Concentration of Stochastic Approximation with Time-Inhomogeneous Markov Chains
by: Singh, Rahul, et al.
Published: (2026) -
Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
by: Kar, Avik, et al.
Published: (2024) -
Enabling Off-Policy Imitation Learning with Deep Actor Critic Stabilization
by: Sen, Sayambhu, et al.
Published: (2025) -
Finite-Time Bounds for Two-Time-Scale Stochastic Approximation with Arbitrary Norm Contractions and Markovian Noise
by: Chandak, Siddharth, et al.
Published: (2025)