Reinforcement Learning for Exponential Utility: Algorithms and Convergence in Discounted MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Thoppe, Gugan, Prashanth, L. A., Naskar, Ankur, Bhat, Sanjay |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Risk Estimation in a Markov Cost Process: Lower and Upper Bounds
by: Thoppe, Gugan, et al.
Published: (2023)
by: Thoppe, Gugan, et al.
Published: (2023)
Parameter-free Optimal Rates for Nonlinear Semi-Norm Contractions with Applications to $Q$-Learning
by: Naskar, Ankur, et al.
Published: (2025)
by: Naskar, Ankur, et al.
Published: (2025)
Parameter-Free Federated TD Learning with Markov Noise in Heterogeneous Environments
by: Naskar, Ankur, et al.
Published: (2025)
by: Naskar, Ankur, et al.
Published: (2025)
Reinforcement Learning with Quasi-Hyperbolic Discounting
by: Eshwar, S. R., et al.
Published: (2024)
by: Eshwar, S. R., et al.
Published: (2024)
Does DQN Learn?
by: Gopalan, Aditya, et al.
Published: (2022)
by: Gopalan, Aditya, et al.
Published: (2022)
Global Convergence Guarantees for Federated Policy Gradient Methods with Adversaries
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs
by: Kash, Ian A., et al.
Published: (2022)
by: Kash, Ian A., et al.
Published: (2022)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Tight Convergence Rates for Online Distributed Linear Estimation with Adversarial Measurements
by: Roy, Nibedita, et al.
Published: (2026)
by: Roy, Nibedita, et al.
Published: (2026)
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023)
by: Dalal, Gal, et al.
Published: (2023)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
by: Viano, Luca, et al.
Published: (2024)
by: Viano, Luca, et al.
Published: (2024)
Optimization of utility-based shortfall risk: A non-asymptotic viewpoint
by: Gupte, Sumedh, et al.
Published: (2023)
by: Gupte, Sumedh, et al.
Published: (2023)
Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor
by: Grand-Clément, Julien, et al.
Published: (2023)
by: Grand-Clément, Julien, et al.
Published: (2023)
A Finite-Sample Analysis of an Actor-Critic Algorithm for Mean-Variance Optimization in a Discounted MDP
by: Sangadi, Tejaram, et al.
Published: (2024)
by: Sangadi, Tejaram, et al.
Published: (2024)
Efficiently Solving Discounted MDPs with Predictions on Transition Matrices
by: Lyu, Lixing, et al.
Published: (2025)
by: Lyu, Lixing, et al.
Published: (2025)
Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting
by: Skalse, Joar, et al.
Published: (2024)
by: Skalse, Joar, et al.
Published: (2024)
Adversary-Robust Learning from Fully Asynchronous Directional Derivative Estimates
by: Paul, Anik Kumar, et al.
Published: (2026)
by: Paul, Anik Kumar, et al.
Published: (2026)
What Can Be Recovered Under Sparse Adversarial Corruption? Assumption-Free Theory for Linear Measurements
by: Halder, Vishal, et al.
Published: (2025)
by: Halder, Vishal, et al.
Published: (2025)
Reliable Policy Iteration: Performance Robustness Across Architecture and Environment Perturbations
by: Eshwar, S. R., et al.
Published: (2025)
by: Eshwar, S. R., et al.
Published: (2025)
A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Learning to Reason Efficiently with Discounted Reinforcement Learning
by: Ayoub, Alex, et al.
Published: (2025)
by: Ayoub, Alex, et al.
Published: (2025)
On Value Iteration Convergence in Connected MDPs
by: Mustafin, Arsenii, et al.
Published: (2024)
by: Mustafin, Arsenii, et al.
Published: (2024)
Risk-sensitive reinforcement learning using expectiles, shortfall risk and optimized certainty equivalent risk
by: Gupte, Sumedh, et al.
Published: (2026)
by: Gupte, Sumedh, et al.
Published: (2026)
Monotone and Conservative Policy Iteration Beyond the Tabular Case
by: Eshwar, S. R., et al.
Published: (2025)
by: Eshwar, S. R., et al.
Published: (2025)
Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
by: Manivannan, Sanjeev, et al.
Published: (2026)
by: Manivannan, Sanjeev, et al.
Published: (2026)
No-Regret Reinforcement Learning in Smooth MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
by: Zurek, Matthew, et al.
Published: (2024)
by: Zurek, Matthew, et al.
Published: (2024)
Reinforcement Learning from Adversarial Preferences in Tabular MDPs
by: Tsuchiya, Taira, et al.
Published: (2025)
by: Tsuchiya, Taira, et al.
Published: (2025)
Teaching Precommitted Agents: Model-Free Policy Evaluation and Control in Quasi-Hyperbolic Discounted MDPs
by: Eshwar, S. R.
Published: (2025)
by: Eshwar, S. R.
Published: (2025)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Policy Gradient in Robust MDPs with Global Convergence Guarantee
by: Wang, Qiuhao, et al.
Published: (2022)
by: Wang, Qiuhao, et al.
Published: (2022)
Local Linearity: the Key for No-regret Reinforcement Learning in Continuous MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
Decoupling Time and Risk: Risk-Sensitive Reinforcement Learning with General Discounting
by: Moghimi, Mehrdad, et al.
Published: (2026)
by: Moghimi, Mehrdad, et al.
Published: (2026)
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs
by: Lu, Michael, et al.
Published: (2026)
by: Lu, Michael, et al.
Published: (2026)
Recursive Entropic Risk Optimization in Discounted MDPs: Sample Complexity Bounds with a Generative Model
by: Mortensen, Oliver, et al.
Published: (2025)
by: Mortensen, Oliver, et al.
Published: (2025)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
by: Zhang, Zhongjun, et al.
Published: (2026)
by: Zhang, Zhongjun, et al.
Published: (2026)
Joint MDPs and Reinforcement Learning in Coupled-Dynamics Environments
by: Kaya, Ege C., et al.
Published: (2026)
by: Kaya, Ege C., et al.
Published: (2026)
Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs
by: Tan, Kevin, et al.
Published: (2024)
by: Tan, Kevin, et al.
Published: (2024)
Stable and Robust Deep Learning By Hyperbolic Tangent Exponential Linear Unit (TeLU)
by: Fernandez, Alfredo, et al.
Published: (2024)
by: Fernandez, Alfredo, et al.
Published: (2024)
Convergent Reinforcement Learning Algorithms for Stochastic Shortest Path Problem
by: Guin, Soumyajit, et al.
Published: (2025)
by: Guin, Soumyajit, et al.
Published: (2025)
Similar Items
-
Risk Estimation in a Markov Cost Process: Lower and Upper Bounds
by: Thoppe, Gugan, et al.
Published: (2023) -
Parameter-free Optimal Rates for Nonlinear Semi-Norm Contractions with Applications to $Q$-Learning
by: Naskar, Ankur, et al.
Published: (2025) -
Parameter-Free Federated TD Learning with Markov Noise in Heterogeneous Environments
by: Naskar, Ankur, et al.
Published: (2025) -
Reinforcement Learning with Quasi-Hyperbolic Discounting
by: Eshwar, S. R., et al.
Published: (2024) -
Does DQN Learn?
by: Gopalan, Aditya, et al.
Published: (2022)