Bayesian Risk-Sensitive Policy Optimization For MDPs With General Loss Functions
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xiaoshuang, Lin, Yifan, Zhou, Enlu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evolving Robustness--Exploration Trade-off in Online Reinforcement Learning via Quantile Bayesian Risk MDPs
by: Song, Meichen, et al.
Published: (2026)
by: Song, Meichen, et al.
Published: (2026)
Online Bayesian Risk-Averse Reinforcement Learning
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate
by: Lin, Yifan, et al.
Published: (2024)
by: Lin, Yifan, et al.
Published: (2024)
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
by: Zhang, Runyu, et al.
Published: (2023)
by: Zhang, Runyu, et al.
Published: (2023)
Landscape of Policy Optimization for Finite Horizon MDPs with General State and Action
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
Adaptive Simulation Experiment for LLM Policy Optimization
by: Hu, Mingjie, et al.
Published: (2026)
by: Hu, Mingjie, et al.
Published: (2026)
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
by: Levy, Orin, et al.
Published: (2026)
by: Levy, Orin, et al.
Published: (2026)
Approximate Bilevel Difference Convex Programming for Bayesian Risk Markov Decision Processes
by: Lin, Yifan, et al.
Published: (2023)
by: Lin, Yifan, et al.
Published: (2023)
Ranking and Selection with Simultaneous Input Data Collection
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
by: Zhong, Han, et al.
Published: (2021)
by: Zhong, Han, et al.
Published: (2021)
Curiosity is Knowledge: Self-Consistent Learning and No-Regret Optimization with Active Inference
by: Li, Yingke, et al.
Published: (2026)
by: Li, Yingke, et al.
Published: (2026)
Pragmatic Curiosity: A Unified Framework for Hybrid Learning and Optimization via Active Inference
by: Li, Yingke, et al.
Published: (2026)
by: Li, Yingke, et al.
Published: (2026)
Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization
by: Tiapkin, Daniil, et al.
Published: (2024)
by: Tiapkin, Daniil, et al.
Published: (2024)
Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
by: Stradi, Francesco Emanuele, et al.
Published: (2024)
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
by: Agnihotri, Akhil, et al.
Published: (2023)
by: Agnihotri, Akhil, et al.
Published: (2023)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
by: Schlisselberg, Ofir, et al.
Published: (2026)
by: Schlisselberg, Ofir, et al.
Published: (2026)
Policy Gradient in Robust MDPs with Global Convergence Guarantee
by: Wang, Qiuhao, et al.
Published: (2022)
by: Wang, Qiuhao, et al.
Published: (2022)
Planning and Learning in Average Risk-aware MDPs
by: Wang, Weikai, et al.
Published: (2025)
by: Wang, Weikai, et al.
Published: (2025)
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
by: Mondal, Washim Uddin, et al.
Published: (2024)
by: Mondal, Washim Uddin, et al.
Published: (2024)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
by: Lancewicki, Tal, et al.
Published: (2025)
by: Lancewicki, Tal, et al.
Published: (2025)
Constraint-Generation Policy Optimization (CGPO): Nonlinear Programming for Policy Optimization in Mixed Discrete-Continuous MDPs
by: Gimelfarb, Michael, et al.
Published: (2024)
by: Gimelfarb, Michael, et al.
Published: (2024)
RASR: Risk-Averse Soft-Robust MDPs with EVaR and Entropic Risk
by: Hau, Jia Lin, et al.
Published: (2022)
by: Hau, Jia Lin, et al.
Published: (2022)
Provable Risk-Sensitive Distributional Reinforcement Learning with General Function Approximation
by: Chen, Yu, et al.
Published: (2024)
by: Chen, Yu, et al.
Published: (2024)
Confident Natural Policy Gradient for Local Planning in $q_π$-realizable Constrained MDPs
by: Tian, Tian, et al.
Published: (2024)
by: Tian, Tian, et al.
Published: (2024)
Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
by: Ito, Shinji, et al.
Published: (2025)
by: Ito, Shinji, et al.
Published: (2025)
Global Convergence of Average Reward Constrained MDPs with Neural Critic and General Policy Parameterization
by: Satheesh, Anirudh, et al.
Published: (2026)
by: Satheesh, Anirudh, et al.
Published: (2026)
Recursive Entropic Risk Optimization in Discounted MDPs: Sample Complexity Bounds with a Generative Model
by: Mortensen, Oliver, et al.
Published: (2025)
by: Mortensen, Oliver, et al.
Published: (2025)
Generative Bayesian Optimization: Generative Models as Acquisition Functions
by: Oliveira, Rafael, et al.
Published: (2025)
by: Oliveira, Rafael, et al.
Published: (2025)
Bayesian Risk-averse Model Predictive Control with Consistency and Stability Guarantees
by: Li, Yingke, et al.
Published: (2025)
by: Li, Yingke, et al.
Published: (2025)
A Unified Framework for Tabular Generative Modeling: Loss Functions, Benchmarks, and Improved Multi-objective Bayesian Optimization Approaches
by: Vu, Minh H., et al.
Published: (2024)
by: Vu, Minh H., et al.
Published: (2024)
Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs
by: Lu, Michael, et al.
Published: (2024)
by: Lu, Michael, et al.
Published: (2024)
Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief
by: Lin, Hongqiang, et al.
Published: (2026)
by: Lin, Hongqiang, et al.
Published: (2026)
Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
by: Liu, Xingtu, et al.
Published: (2025)
by: Liu, Xingtu, et al.
Published: (2025)
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
Sample-efficient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation
by: He, Jianliang, et al.
Published: (2024)
by: He, Jianliang, et al.
Published: (2024)
Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs
by: Lee, Jongmin, et al.
Published: (2025)
by: Lee, Jongmin, et al.
Published: (2025)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
by: Wang, Shengbo, et al.
Published: (2026)
by: Wang, Shengbo, et al.
Published: (2026)
Is Pure Exploitation Sufficient in Exogenous MDPs with Linear Function Approximation?
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
Revisiting Weighted Strategy for Non-stationary Parametric Bandits and MDPs
by: Wang, Jing, et al.
Published: (2026)
by: Wang, Jing, et al.
Published: (2026)
Convergence of Natural Policy Gradient for a Family of Infinite-State Queueing MDPs
by: Grosof, Isaac, et al.
Published: (2024)
by: Grosof, Isaac, et al.
Published: (2024)
Similar Items
-
Evolving Robustness--Exploration Trade-off in Online Reinforcement Learning via Quantile Bayesian Risk MDPs
by: Song, Meichen, et al.
Published: (2026) -
Online Bayesian Risk-Averse Reinforcement Learning
by: Wang, Yuhao, et al.
Published: (2025) -
Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate
by: Lin, Yifan, et al.
Published: (2024) -
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
by: Zhang, Runyu, et al.
Published: (2023) -
Landscape of Policy Optimization for Finite Horizon MDPs with General State and Action
by: Chen, Xin, et al.
Published: (2024)