Robust Parameter Learning for Uncertain MDPs
Fuente:
arXiv
Guardado en:
| Autores principales: | Schnitzer, Yannik, Abate, Alessandro, Parker, David |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Efficient Solution and Learning of Robust Factored MDPs
por: Schnitzer, Yannik, et al.
Publicado: (2025)
por: Schnitzer, Yannik, et al.
Publicado: (2025)
Certifiably Robust Policies for Uncertain Parametric Environments
por: Schnitzer, Yannik, et al.
Publicado: (2024)
por: Schnitzer, Yannik, et al.
Publicado: (2024)
Probabilistic Performance Guarantees for Multi-Task Reinforcement Learning
por: Schnitzer, Yannik, et al.
Publicado: (2026)
por: Schnitzer, Yannik, et al.
Publicado: (2026)
Bisimulation Learning
por: Abate, Alessandro, et al.
Publicado: (2024)
por: Abate, Alessandro, et al.
Publicado: (2024)
Certified Approximate Reachability (CARe): Formal Error Bounds on Deep Learning of Reachable Sets
por: Solanki, Prashant, et al.
Publicado: (2025)
por: Solanki, Prashant, et al.
Publicado: (2025)
Branching Bisimulation Learning
por: Abate, Alessandro, et al.
Publicado: (2025)
por: Abate, Alessandro, et al.
Publicado: (2025)
Time-Constrained Robust MDPs
por: Zouitine, Adil, et al.
Publicado: (2024)
por: Zouitine, Adil, et al.
Publicado: (2024)
Quantifying the Sensitivity of Inverse Reinforcement Learning to Misspecification
por: Skalse, Joar, et al.
Publicado: (2024)
por: Skalse, Joar, et al.
Publicado: (2024)
Ensuring Safety in an Uncertain Environment: Constrained MDPs via Stochastic Thresholds
por: Zuo, Qian, et al.
Publicado: (2025)
por: Zuo, Qian, et al.
Publicado: (2025)
Partial Identifiability and Misspecification in Inverse Reinforcement Learning
por: Skalse, Joar, et al.
Publicado: (2024)
por: Skalse, Joar, et al.
Publicado: (2024)
Partial Identifiability in Inverse Reinforcement Learning For Agents With Non-Exponential Discounting
por: Skalse, Joar, et al.
Publicado: (2024)
por: Skalse, Joar, et al.
Publicado: (2024)
DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RL
por: Jackermeier, Mathias, et al.
Publicado: (2024)
por: Jackermeier, Mathias, et al.
Publicado: (2024)
Solving Robust MDPs through No-Regret Dynamics
por: Guha, Etash Kumar
Publicado: (2023)
por: Guha, Etash Kumar
Publicado: (2023)
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
por: Zhang, Runyu, et al.
Publicado: (2023)
por: Zhang, Runyu, et al.
Publicado: (2023)
Neural Proofs for Sound Verification and Control of Complex Systems
por: Abate, Alessandro
Publicado: (2025)
por: Abate, Alessandro
Publicado: (2025)
Policy Gradient in Robust MDPs with Global Convergence Guarantee
por: Wang, Qiuhao, et al.
Publicado: (2022)
por: Wang, Qiuhao, et al.
Publicado: (2022)
On the Limitations of Markovian Rewards to Express Multi-Objective, Risk-Sensitive, and Modal Tasks
por: Skalse, Joar, et al.
Publicado: (2024)
por: Skalse, Joar, et al.
Publicado: (2024)
Multi-Property Synthesis
por: Weinhuber, Christoph, et al.
Publicado: (2026)
por: Weinhuber, Christoph, et al.
Publicado: (2026)
Temporal-Difference Variational Continual Learning
por: Melo, Luckeciano C., et al.
Publicado: (2024)
por: Melo, Luckeciano C., et al.
Publicado: (2024)
Truly No-Regret Learning in Constrained MDPs
por: Müller, Adrian, et al.
Publicado: (2024)
por: Müller, Adrian, et al.
Publicado: (2024)
PlatoLTL: Learning to Generalize Across Symbols in LTL Instructions for Multi-Task RL
por: Cloete, Jacques, et al.
Publicado: (2026)
por: Cloete, Jacques, et al.
Publicado: (2026)
Walking the Values in Bayesian Inverse Reinforcement Learning
por: Bajgar, Ondrej, et al.
Publicado: (2024)
por: Bajgar, Ondrej, et al.
Publicado: (2024)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
por: Melo, Luckeciano C., et al.
Publicado: (2025)
por: Melo, Luckeciano C., et al.
Publicado: (2025)
SPoRt -- Safe Policy Ratio: Certified Training and Deployment of Task Policies in Model-Free RL
por: Cloete, Jacques, et al.
Publicado: (2025)
por: Cloete, Jacques, et al.
Publicado: (2025)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
por: Hong, Kihyuk, et al.
Publicado: (2024)
por: Hong, Kihyuk, et al.
Publicado: (2024)
Robust Shielding for Safe Reinforcement Learning
por: Court, Edwin Hamel-De le, et al.
Publicado: (2026)
por: Court, Edwin Hamel-De le, et al.
Publicado: (2026)
Efficient Duple Perturbation Robustness in Low-rank MDPs
por: Hu, Yang, et al.
Publicado: (2024)
por: Hu, Yang, et al.
Publicado: (2024)
Provably Efficient Algorithms for S- and Non-Rectangular Robust MDPs with General Parameterization
por: Satheesh, Anirudh, et al.
Publicado: (2026)
por: Satheesh, Anirudh, et al.
Publicado: (2026)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
por: Gadot, Uri, et al.
Publicado: (2023)
por: Gadot, Uri, et al.
Publicado: (2023)
Learning Adversarial MDPs with Stochastic Hard Constraints
por: Stradi, Francesco Emanuele, et al.
Publicado: (2024)
por: Stradi, Francesco Emanuele, et al.
Publicado: (2024)
Evolving Robustness--Exploration Trade-off in Online Reinforcement Learning via Quantile Bayesian Risk MDPs
por: Song, Meichen, et al.
Publicado: (2026)
por: Song, Meichen, et al.
Publicado: (2026)
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
por: Wang, Kaixin, et al.
Publicado: (2023)
por: Wang, Kaixin, et al.
Publicado: (2023)
No-Regret Reinforcement Learning in Smooth MDPs
por: Maran, Davide, et al.
Publicado: (2024)
por: Maran, Davide, et al.
Publicado: (2024)
Primal-Only Actor Critic Algorithm for Robust Constrained Average Cost MDPs
por: Satheesh, Anirudh, et al.
Publicado: (2025)
por: Satheesh, Anirudh, et al.
Publicado: (2025)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
por: Schlisselberg, Ofir, et al.
Publicado: (2026)
por: Schlisselberg, Ofir, et al.
Publicado: (2026)
Reinforcement Learning from Adversarial Preferences in Tabular MDPs
por: Tsuchiya, Taira, et al.
Publicado: (2025)
por: Tsuchiya, Taira, et al.
Publicado: (2025)
Offline Oracle-Efficient Learning for Contextual MDPs via Layerwise Exploration-Exploitation Tradeoff
por: Qian, Jian, et al.
Publicado: (2024)
por: Qian, Jian, et al.
Publicado: (2024)
Evaluating and Learning Robust Bandit Policies Under Uncertain Causal Mechanisms
por: Avery, Katherine, et al.
Publicado: (2025)
por: Avery, Katherine, et al.
Publicado: (2025)
Zero-Shot Instruction Following in RL via Structured LTL Representations
por: Giuri, Mattia, et al.
Publicado: (2025)
por: Giuri, Mattia, et al.
Publicado: (2025)
Learning Robust Policies for Uncertain Parametric Markov Decision Processes
por: Rickard, Luke, et al.
Publicado: (2023)
por: Rickard, Luke, et al.
Publicado: (2023)
Ejemplares similares
-
Efficient Solution and Learning of Robust Factored MDPs
por: Schnitzer, Yannik, et al.
Publicado: (2025) -
Certifiably Robust Policies for Uncertain Parametric Environments
por: Schnitzer, Yannik, et al.
Publicado: (2024) -
Probabilistic Performance Guarantees for Multi-Task Reinforcement Learning
por: Schnitzer, Yannik, et al.
Publicado: (2026) -
Bisimulation Learning
por: Abate, Alessandro, et al.
Publicado: (2024) -
Certified Approximate Reachability (CARe): Formal Error Bounds on Deep Learning of Reachable Sets
por: Solanki, Prashant, et al.
Publicado: (2025)