Saved in:
| Main Authors: | Aberdeen, Douglas, Baxter, Jonathan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.03204 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Recurrent Natural Policy Gradient for POMDPs
by: Cayci, Semih, et al.
Published: (2024)
by: Cayci, Semih, et al.
Published: (2024)
A Multi-Agent, Policy-Gradient approach to Network Routing
by: Tao, Nigel, et al.
Published: (2025)
by: Tao, Nigel, et al.
Published: (2025)
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2025)
by: Galesloot, Maris F. L., et al.
Published: (2025)
Reinforcement Learning in POMDP's via Direct Gradient Ascent
by: Baxter, Jonathan, et al.
Published: (2025)
by: Baxter, Jonathan, et al.
Published: (2025)
Memoryless Policy Iteration for Episodic POMDPs
by: van Zuijlen, Roy, et al.
Published: (2025)
by: van Zuijlen, Roy, et al.
Published: (2025)
Internal State-Based Policy Gradient Methods for Partially Observable Markov Potential Games
by: Yang, Wonseok, et al.
Published: (2026)
by: Yang, Wonseok, et al.
Published: (2026)
Reinforcement Learning From State and Temporal Differences
by: Weaver, Lex, et al.
Published: (2025)
by: Weaver, Lex, et al.
Published: (2025)
Scalable Policy-Based RL Algorithms for POMDPs
by: Anjarlekar, Ameya, et al.
Published: (2025)
by: Anjarlekar, Ameya, et al.
Published: (2025)
Online Planning in POMDPs with State-Requests
by: Avalos, Raphael, et al.
Published: (2024)
by: Avalos, Raphael, et al.
Published: (2024)
Sequential Monte Carlo for Policy Optimization in Continuous POMDPs
by: Abdulsamad, Hany, et al.
Published: (2025)
by: Abdulsamad, Hany, et al.
Published: (2025)
Model-Based Learning of Near-Optimal Finite-Window Policies in POMDPs
by: Jordan, Philip, et al.
Published: (2026)
by: Jordan, Philip, et al.
Published: (2026)
The Evolution of Learning Algorithms for Artificial Neural Networks
by: Baxter, Jonathan
Published: (2025)
by: Baxter, Jonathan
Published: (2025)
Policy Gradient Methods in the Presence of Symmetries and State Abstractions
by: Panangaden, Prakash, et al.
Published: (2023)
by: Panangaden, Prakash, et al.
Published: (2023)
How to Explore with Belief: State Entropy Maximization in POMDPs
by: Zamboni, Riccardo, et al.
Published: (2024)
by: Zamboni, Riccardo, et al.
Published: (2024)
Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPs
by: Zhang, Yuheng, et al.
Published: (2025)
by: Zhang, Yuheng, et al.
Published: (2025)
A result relating convex n-widths to covering numbers with some applications to neural networks
by: Baxter, Jonathan, et al.
Published: (2025)
by: Baxter, Jonathan, et al.
Published: (2025)
Learning General Policies with Policy Gradient Methods
by: Ståhlberg, Simon, et al.
Published: (2025)
by: Ståhlberg, Simon, et al.
Published: (2025)
Rethinking Transformers in Solving POMDPs
by: Lu, Chenhao, et al.
Published: (2024)
by: Lu, Chenhao, et al.
Published: (2024)
When Do Off-Policy and On-Policy Policy Gradient Methods Align?
by: Mambelli, Davide, et al.
Published: (2024)
by: Mambelli, Davide, et al.
Published: (2024)
Explainable Representation of Finite-Memory Policies for POMDPs using Decision Trees
by: Azeem, Muqsit, et al.
Published: (2024)
by: Azeem, Muqsit, et al.
Published: (2024)
Learning Interpretable Policies in Hindsight-Observable POMDPs through Partially Supervised Reinforcement Learning
by: Lanier, Michael, et al.
Published: (2024)
by: Lanier, Michael, et al.
Published: (2024)
Toward Learning POMDPs Beyond Full-Rank Actions and State Observability
by: Shaw, Seiji, et al.
Published: (2026)
by: Shaw, Seiji, et al.
Published: (2026)
Policy Gradient Methods for Distortion Risk Measures
by: Vijayan, Nithia, et al.
Published: (2021)
by: Vijayan, Nithia, et al.
Published: (2021)
Perception-Based Beliefs for POMDPs with Visual Observations
by: Schäfers, Miriam, et al.
Published: (2026)
by: Schäfers, Miriam, et al.
Published: (2026)
Elementary Analysis of Policy Gradient Methods
by: Liu, Jiacai, et al.
Published: (2024)
by: Liu, Jiacai, et al.
Published: (2024)
Learning Logic Specifications for Policy Guidance in POMDPs: an Inductive Logic Programming Approach
by: Meli, Daniele, et al.
Published: (2024)
by: Meli, Daniele, et al.
Published: (2024)
Gradient Methods with Online Scaling
by: Gao, Wenzhi, et al.
Published: (2024)
by: Gao, Wenzhi, et al.
Published: (2024)
Reevaluating Policy Gradient Methods for Imperfect-Information Games
by: Rudolph, Max, et al.
Published: (2025)
by: Rudolph, Max, et al.
Published: (2025)
Mollification Effects of Policy Gradient Methods
by: Wang, Tao, et al.
Published: (2024)
by: Wang, Tao, et al.
Published: (2024)
Periodic agent-state based Q-learning for POMDPs
by: Sinha, Amit, et al.
Published: (2024)
by: Sinha, Amit, et al.
Published: (2024)
Logit Dynamics in Softmax Policy Gradient Methods
by: Li, Yingru
Published: (2025)
by: Li, Yingru
Published: (2025)
Pessimistic Iterative Planning with RNNs for Robust POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2024)
by: Galesloot, Maris F. L., et al.
Published: (2024)
Approximate Control for Continuous-Time POMDPs
by: Eich, Yannick, et al.
Published: (2024)
by: Eich, Yannick, et al.
Published: (2024)
Ordering-based Conditions for Global Convergence of Policy Gradient Methods
by: Mei, Jincheng, et al.
Published: (2025)
by: Mei, Jincheng, et al.
Published: (2025)
Convergence of regularized agent-state-based Q-learning in POMDPs
by: Sinha, Amit, et al.
Published: (2025)
by: Sinha, Amit, et al.
Published: (2025)
The Limits of Pure Exploration in POMDPs: When the Observation Entropy is Enough
by: Zamboni, Riccardo, et al.
Published: (2024)
by: Zamboni, Riccardo, et al.
Published: (2024)
Beyond Exact Gradients: Convergence of Stochastic Soft-Max Policy Gradient Methods with Entropy Regularization
by: Ding, Yuhao, et al.
Published: (2021)
by: Ding, Yuhao, et al.
Published: (2021)
Policy Gradient Methods for Non-Markovian Reinforcement Learning
by: Kar, Avik, et al.
Published: (2026)
by: Kar, Avik, et al.
Published: (2026)
Matrix Low-Rank Approximation For Policy Gradient Methods
by: Rozada, Sergio, et al.
Published: (2024)
by: Rozada, Sergio, et al.
Published: (2024)
Missingness-MDPs: Bridging the Theory of Missing Data and POMDPs
by: Wendland, Joshua, et al.
Published: (2026)
by: Wendland, Joshua, et al.
Published: (2026)
Similar Items
-
Recurrent Natural Policy Gradient for POMDPs
by: Cayci, Semih, et al.
Published: (2024) -
A Multi-Agent, Policy-Gradient approach to Network Routing
by: Tao, Nigel, et al.
Published: (2025) -
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2025) -
Reinforcement Learning in POMDP's via Direct Gradient Ascent
by: Baxter, Jonathan, et al.
Published: (2025) -
Memoryless Policy Iteration for Episodic POMDPs
by: van Zuijlen, Roy, et al.
Published: (2025)