Howard's Policy Iteration is Subexponential for Deterministic Markov Decision Problems with Rewards of Fixed Bit-size and Arbitrary Discount Factor
Fuente:
arXiv
Saved in:
| Main Authors: | Mukherjee, Dibyangshu, Kalyanakrishnan, Shivaram |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Computation of Blackwell Optimal Policies using Rational Functions
by: Mukherjee, Dibyangshu, et al.
Published: (2025)
by: Mukherjee, Dibyangshu, et al.
Published: (2025)
Lower Bound on Howard Policy Iteration for Deterministic Markov Decision Processes
by: Asadi, Ali, et al.
Published: (2025)
by: Asadi, Ali, et al.
Published: (2025)
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
by: Shah, Anvay, et al.
Published: (2026)
by: Shah, Anvay, et al.
Published: (2026)
Improved Sample Complexity Analysis of Natural Policy Gradient Algorithm with General Parameterization for Infinite Horizon Discounted Reward Markov Decision Processes
by: Mondal, Washim Uddin, et al.
Published: (2023)
by: Mondal, Washim Uddin, et al.
Published: (2023)
Value Iteration with Guessing for Markov Chains and Markov Decision Processes
by: Chatterjee, Krishnendu, et al.
Published: (2025)
by: Chatterjee, Krishnendu, et al.
Published: (2025)
Regret Analysis of Policy Gradient Algorithm for Infinite Horizon Average Reward Markov Decision Processes
by: Bai, Qinbo, et al.
Published: (2023)
by: Bai, Qinbo, et al.
Published: (2023)
On the Convergence of Modified Policy Iteration in Risk Sensitive Exponential Cost Markov Decision Processes
by: Murthy, Yashaswini, et al.
Published: (2023)
by: Murthy, Yashaswini, et al.
Published: (2023)
Optimal Decision Tree Policies for Markov Decision Processes
by: Vos, Daniël, et al.
Published: (2023)
by: Vos, Daniël, et al.
Published: (2023)
Hierarchical Average-Reward Linearly-solvable Markov Decision Processes
by: Infante, Guillermo, et al.
Published: (2024)
by: Infante, Guillermo, et al.
Published: (2024)
Robust Reward Design for Markov Decision Processes
by: Wu, Shuo, et al.
Published: (2024)
by: Wu, Shuo, et al.
Published: (2024)
Policy Gradient for Robust Markov Decision Processes
by: Wang, Qiuhao, et al.
Published: (2024)
by: Wang, Qiuhao, et al.
Published: (2024)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
by: Singha, Disha
Published: (2026)
by: Singha, Disha
Published: (2026)
Offline Policy Evaluation for Manipulation Policies via Discounted Liveness Formulation
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Supply Chain Optimization via Generative Simulation and Iterative Decision Policies
by: Bai, Haoyue, et al.
Published: (2025)
by: Bai, Haoyue, et al.
Published: (2025)
Provably Efficient Reward Transfer in Reinforcement Learning with Discrete Markov Decision Processes
by: Vora, Kevin, et al.
Published: (2025)
by: Vora, Kevin, et al.
Published: (2025)
Best-Effort Policies for Robust Markov Decision Processes
by: Abate, Alessandro, et al.
Published: (2025)
by: Abate, Alessandro, et al.
Published: (2025)
SPOT: Scalable Policy Optimization with Trees for Markov Decision Processes
by: Xiong, Xuyuan, et al.
Published: (2025)
by: Xiong, Xuyuan, et al.
Published: (2025)
Enhancing Table Reasoning with Deterministic Table-State Rewards
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
by: Kwok, Tung Sum Thomas, et al.
Published: (2026)
Conformal Off-Policy Evaluation in Markov Decision Processes
by: Foffano, Daniele, et al.
Published: (2023)
by: Foffano, Daniele, et al.
Published: (2023)
Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes
by: Blaser, Ethan, et al.
Published: (2026)
by: Blaser, Ethan, et al.
Published: (2026)
Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision Processes
by: Bennett, Andrew, et al.
Published: (2024)
by: Bennett, Andrew, et al.
Published: (2024)
Beyond the Bellman Fixed Point: Geometry and Fast Policy Identification in Value Iteration
by: Lee, Donghwan
Published: (2026)
by: Lee, Donghwan
Published: (2026)
Quantile Markov Decision Process
by: Li, Xiaocheng, et al.
Published: (2017)
by: Li, Xiaocheng, et al.
Published: (2017)
Creativity and Markov Decision Processes
by: Lahikainen, Joonas, et al.
Published: (2024)
by: Lahikainen, Joonas, et al.
Published: (2024)
GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets
by: Yao, Zhangyang, et al.
Published: (2026)
by: Yao, Zhangyang, et al.
Published: (2026)
LLMs for High-Frequency Decision-Making: Normalized Action Reward-Guided Consistency Policy Optimization
by: Zhao, Yang, et al.
Published: (2026)
by: Zhao, Yang, et al.
Published: (2026)
Inferring Reward Machines and Transition Machines from Partially Observable Markov Decision Processes
by: Wu, Yuly, et al.
Published: (2025)
by: Wu, Yuly, et al.
Published: (2025)
Quantum Speedups in Regret Analysis of Infinite Horizon Average-Reward Markov Decision Processes
by: Ganguly, Bhargav, et al.
Published: (2023)
by: Ganguly, Bhargav, et al.
Published: (2023)
Policy Regularized Distributionally Robust Markov Decision Processes with Linear Function Approximation
by: Gu, Jingwen, et al.
Published: (2025)
by: Gu, Jingwen, et al.
Published: (2025)
Counterfactual Influence in Markov Decision Processes
by: Kazemi, Milad, et al.
Published: (2024)
by: Kazemi, Milad, et al.
Published: (2024)
Reliable Policy Iteration: Performance Robustness Across Architecture and Environment Perturbations
by: Eshwar, S. R., et al.
Published: (2025)
by: Eshwar, S. R., et al.
Published: (2025)
Burning RED: Unlocking Subtask-Driven Reinforcement Learning and Risk-Awareness in Average-Reward Markov Decision Processes
by: Rojas, Juan Sebastian, et al.
Published: (2024)
by: Rojas, Juan Sebastian, et al.
Published: (2024)
Tesserae: Scalable Placement Policies for Deep Learning Workloads
by: Bian, Song, et al.
Published: (2025)
by: Bian, Song, et al.
Published: (2025)
Detecting Hidden Triggers: Mapping Non-Markov Reward Functions to Markov
by: Hyde, Gregory, et al.
Published: (2024)
by: Hyde, Gregory, et al.
Published: (2024)
Policy Gradient Algorithms with Monte Carlo Tree Learning for Non-Markov Decision Processes
by: Morimura, Tetsuro, et al.
Published: (2022)
by: Morimura, Tetsuro, et al.
Published: (2022)
Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
by: Manivannan, Sanjeev, et al.
Published: (2026)
by: Manivannan, Sanjeev, et al.
Published: (2026)
Sharpe Ratio Optimization in Markov Decision Processes
by: Ma, Shuai, et al.
Published: (2025)
by: Ma, Shuai, et al.
Published: (2025)
Robust Counterfactual Inference in Markov Decision Processes
by: Lally, Jessica, et al.
Published: (2025)
by: Lally, Jessica, et al.
Published: (2025)
Spontaneous Reward Hacking in Iterative Self-Refinement
by: Pan, Jane, et al.
Published: (2024)
by: Pan, Jane, et al.
Published: (2024)
A View of the Certainty-Equivalence Method for PAC RL as an Application of the Trajectory Tree Method
by: Kalyanakrishnan, Shivaram, et al.
Published: (2025)
by: Kalyanakrishnan, Shivaram, et al.
Published: (2025)
Similar Items
-
Efficient Computation of Blackwell Optimal Policies using Rational Functions
by: Mukherjee, Dibyangshu, et al.
Published: (2025) -
Lower Bound on Howard Policy Iteration for Deterministic Markov Decision Processes
by: Asadi, Ali, et al.
Published: (2025) -
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
by: Shah, Anvay, et al.
Published: (2026) -
Improved Sample Complexity Analysis of Natural Policy Gradient Algorithm with General Parameterization for Infinite Horizon Discounted Reward Markov Decision Processes
by: Mondal, Washim Uddin, et al.
Published: (2023) -
Value Iteration with Guessing for Markov Chains and Markov Decision Processes
by: Chatterjee, Krishnendu, et al.
Published: (2025)