Scalable Policy-Based RL Algorithms for POMDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Anjarlekar, Ameya, Etesami, Rasoul, Srikant, R |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Striking a Balance: An Optimal Mechanism Design for Heterogenous Differentially Private Data Acquisition for Logistic Regression
by: Anjarlekar, Ameya, et al.
Published: (2023)
by: Anjarlekar, Ameya, et al.
Published: (2023)
Decentralized and Uncoordinated Learning of Stable Matchings: A Game-Theoretic Approach
by: Etesami, S. Rasoul, et al.
Published: (2024)
by: Etesami, S. Rasoul, et al.
Published: (2024)
LLM Unlearning using Gradient Ratio-Based Influence Estimation and Noise Injection
by: Anjarlekar, Ameya, et al.
Published: (2025)
by: Anjarlekar, Ameya, et al.
Published: (2025)
FedGTST: Boosting Global Transferability of Federated Models via Statistics Tuning
by: Ma, Evelyn, et al.
Published: (2024)
by: Ma, Evelyn, et al.
Published: (2024)
A Review of Online Diffusion Policy RL Algorithms for Scalable Robotic Control
by: Choi, Wonhyeok, et al.
Published: (2026)
by: Choi, Wonhyeok, et al.
Published: (2026)
Sequential Monte Carlo for Policy Optimization in Continuous POMDPs
by: Abdulsamad, Hany, et al.
Published: (2025)
by: Abdulsamad, Hany, et al.
Published: (2025)
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2025)
by: Galesloot, Maris F. L., et al.
Published: (2025)
Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPs
by: Zhang, Yuheng, et al.
Published: (2025)
by: Zhang, Yuheng, et al.
Published: (2025)
On the Convergence of Modified Policy Iteration in Risk Sensitive Exponential Cost Markov Decision Processes
by: Murthy, Yashaswini, et al.
Published: (2023)
by: Murthy, Yashaswini, et al.
Published: (2023)
CaRL: Learning Scalable Planning Policies with Simple Rewards
by: Jaeger, Bernhard, et al.
Published: (2025)
by: Jaeger, Bernhard, et al.
Published: (2025)
Rethinking Transformers in Solving POMDPs
by: Lu, Chenhao, et al.
Published: (2024)
by: Lu, Chenhao, et al.
Published: (2024)
Learning Interpretable Policies in Hindsight-Observable POMDPs through Partially Supervised Reinforcement Learning
by: Lanier, Michael, et al.
Published: (2024)
by: Lanier, Michael, et al.
Published: (2024)
Scalable Offline Model-Based RL with Action Chunks
by: Park, Kwanyoung, et al.
Published: (2025)
by: Park, Kwanyoung, et al.
Published: (2025)
Optimal Experiments for Partial Causal Effect Identification
by: Maringgele, Tobias, et al.
Published: (2026)
by: Maringgele, Tobias, et al.
Published: (2026)
Online Planning in POMDPs with State-Requests
by: Avalos, Raphael, et al.
Published: (2024)
by: Avalos, Raphael, et al.
Published: (2024)
GUARD: Guided Unlearning and Retention via Data Attribution for Large Language Models
by: Niu, Peizhi, et al.
Published: (2025)
by: Niu, Peizhi, et al.
Published: (2025)
Pessimistic Iterative Planning with RNNs for Robust POMDPs
by: Galesloot, Maris F. L., et al.
Published: (2024)
by: Galesloot, Maris F. L., et al.
Published: (2024)
Horizon Reduction Makes RL Scalable
by: Park, Seohong, et al.
Published: (2025)
by: Park, Seohong, et al.
Published: (2025)
Search-contempt: a hybrid MCTS algorithm for training AlphaZero-like engines with better computational efficiency
by: Joshi, Ameya
Published: (2025)
by: Joshi, Ameya
Published: (2025)
Explainable Representation of Finite-Memory Policies for POMDPs using Decision Trees
by: Azeem, Muqsit, et al.
Published: (2024)
by: Azeem, Muqsit, et al.
Published: (2024)
Missingness-MDPs: Bridging the Theory of Missing Data and POMDPs
by: Wendland, Joshua, et al.
Published: (2026)
by: Wendland, Joshua, et al.
Published: (2026)
Value of Information and Reward Specification in Active Inference and POMDPs
by: Wei, Ran
Published: (2024)
by: Wei, Ran
Published: (2024)
How to Explore with Belief: State Entropy Maximization in POMDPs
by: Zamboni, Riccardo, et al.
Published: (2024)
by: Zamboni, Riccardo, et al.
Published: (2024)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
by: Du, Yihan, et al.
Published: (2024)
by: Du, Yihan, et al.
Published: (2024)
Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone
by: Mark, Max Sobol, et al.
Published: (2024)
by: Mark, Max Sobol, et al.
Published: (2024)
Learning Logic Specifications for Policy Guidance in POMDPs: an Inductive Logic Programming Approach
by: Meli, Daniele, et al.
Published: (2024)
by: Meli, Daniele, et al.
Published: (2024)
Partial Policy Gradients for RL in LLMs
by: Mathur, Puneet, et al.
Published: (2026)
by: Mathur, Puneet, et al.
Published: (2026)
On Entropy Control in LLM-RL Algorithms
by: Shen, Han
Published: (2025)
by: Shen, Han
Published: (2025)
Deep RL With Information Constrained Policies: Generalization in Continuous Control
by: Malloy, Tailia, et al.
Published: (2020)
by: Malloy, Tailia, et al.
Published: (2020)
Soft Policy Optimization: Online Off-Policy RL for Sequence Models
by: Cohen, Taco, et al.
Published: (2025)
by: Cohen, Taco, et al.
Published: (2025)
Accelerating Goal-Conditioned RL Algorithms and Research
by: Bortkiewicz, Michał, et al.
Published: (2024)
by: Bortkiewicz, Michał, et al.
Published: (2024)
Generalizing Multi-Step Inverse Models for Representation Learning to Finite-Memory POMDPs
by: Wu, Lili, et al.
Published: (2024)
by: Wu, Lili, et al.
Published: (2024)
A Theoretical Analysis of Soft-Label vs Hard-Label Training in Neural Networks
by: Mandal, Saptarshi, et al.
Published: (2024)
by: Mandal, Saptarshi, et al.
Published: (2024)
Nonasymptotic CLT and Error Bounds for Two-Time-Scale Stochastic Approximation
by: Kong, Seo Taek, et al.
Published: (2025)
by: Kong, Seo Taek, et al.
Published: (2025)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
by: Fakoor, Rasool, et al.
Published: (2026)
by: Fakoor, Rasool, et al.
Published: (2026)
Kernel Metric Learning for In-Sample Off-Policy Evaluation of Deterministic RL Policies
by: Lee, Haanvid, et al.
Published: (2024)
by: Lee, Haanvid, et al.
Published: (2024)
Policy Learning for Off-Dynamics RL with Deficient Support
by: Van, Linh Le Pham, et al.
Published: (2024)
by: Van, Linh Le Pham, et al.
Published: (2024)
MindSpeed RL: Distributed Dataflow for Scalable and Efficient RL Training on Ascend NPU Cluster
by: Feng, Laingjun, et al.
Published: (2025)
by: Feng, Laingjun, et al.
Published: (2025)
Sample-efficient and Scalable Exploration in Continuous-Time RL
by: Iten, Klemens, et al.
Published: (2025)
by: Iten, Klemens, et al.
Published: (2025)
METRA: Scalable Unsupervised RL with Metric-Aware Abstraction
by: Park, Seohong, et al.
Published: (2023)
by: Park, Seohong, et al.
Published: (2023)
Similar Items
-
Striking a Balance: An Optimal Mechanism Design for Heterogenous Differentially Private Data Acquisition for Logistic Regression
by: Anjarlekar, Ameya, et al.
Published: (2023) -
Decentralized and Uncoordinated Learning of Stable Matchings: A Game-Theoretic Approach
by: Etesami, S. Rasoul, et al.
Published: (2024) -
LLM Unlearning using Gradient Ratio-Based Influence Estimation and Noise Injection
by: Anjarlekar, Ameya, et al.
Published: (2025) -
FedGTST: Boosting Global Transferability of Federated Models via Statistics Tuning
by: Ma, Evelyn, et al.
Published: (2024) -
A Review of Online Diffusion Policy RL Algorithms for Scalable Robotic Control
by: Choi, Wonhyeok, et al.
Published: (2026)