Solving Robust MDPs through No-Regret Dynamics
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Guha, Etash Kumar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Diminishing Returns of Width for Continual Learning
von: Guha, Etash, et al.
Veröffentlicht: (2024)
von: Guha, Etash, et al.
Veröffentlicht: (2024)
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
Truly No-Regret Learning in Constrained MDPs
von: Müller, Adrian, et al.
Veröffentlicht: (2024)
von: Müller, Adrian, et al.
Veröffentlicht: (2024)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
von: Gadot, Uri, et al.
Veröffentlicht: (2023)
von: Gadot, Uri, et al.
Veröffentlicht: (2023)
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
von: Wang, Kaixin, et al.
Veröffentlicht: (2023)
von: Wang, Kaixin, et al.
Veröffentlicht: (2023)
Eluder-based Regret for Stochastic Contextual MDPs
von: Levy, Orin, et al.
Veröffentlicht: (2022)
von: Levy, Orin, et al.
Veröffentlicht: (2022)
No-Regret Reinforcement Learning in Smooth MDPs
von: Maran, Davide, et al.
Veröffentlicht: (2024)
von: Maran, Davide, et al.
Veröffentlicht: (2024)
Data- and Variance-dependent Regret Bounds for Online Tabular MDPs
von: Li, Mingyi, et al.
Veröffentlicht: (2026)
von: Li, Mingyi, et al.
Veröffentlicht: (2026)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
von: Cassel, Asaf, et al.
Veröffentlicht: (2024)
Conformal Prediction via Regression-as-Classification
von: Guha, Etash, et al.
Veröffentlicht: (2024)
von: Guha, Etash, et al.
Veröffentlicht: (2024)
Exploring and Applying Audio-Based Sentiment Analysis in Music
von: Jhanji, Etash
Veröffentlicht: (2024)
von: Jhanji, Etash
Veröffentlicht: (2024)
Regret Analysis of Unichain Average Reward Constrained MDPs with General Parameterization
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2026)
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2026)
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs
von: Chen, Shulun, et al.
Veröffentlicht: (2025)
von: Chen, Shulun, et al.
Veröffentlicht: (2025)
Optimal Strong Regret and Violation in Constrained MDPs via Policy Optimization
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
Solving Multi-Model MDPs by Coordinate Ascent and Dynamic Programming
von: Su, Xihong, et al.
Veröffentlicht: (2024)
von: Su, Xihong, et al.
Veröffentlicht: (2024)
Regret Analysis of Average-Reward Unichain MDPs via an Actor-Critic Approach
von: Ganesh, Swetha, et al.
Veröffentlicht: (2025)
von: Ganesh, Swetha, et al.
Veröffentlicht: (2025)
Towards Optimal Differentially Private Regret Bounds in Linear MDPs
von: Sahu, Sharan
Veröffentlicht: (2025)
von: Sahu, Sharan
Veröffentlicht: (2025)
Efficient, Low-Regret, Online Reinforcement Learning for Linear MDPs
von: John, Philips George, et al.
Veröffentlicht: (2024)
von: John, Philips George, et al.
Veröffentlicht: (2024)
Efficiently Solving Discounted MDPs with Predictions on Transition Matrices
von: Lyu, Lixing, et al.
Veröffentlicht: (2025)
von: Lyu, Lixing, et al.
Veröffentlicht: (2025)
Time-Constrained Robust MDPs
von: Zouitine, Adil, et al.
Veröffentlicht: (2024)
von: Zouitine, Adil, et al.
Veröffentlicht: (2024)
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
von: Levy, Orin, et al.
Veröffentlicht: (2026)
von: Levy, Orin, et al.
Veröffentlicht: (2026)
Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback
von: Lancewicki, Tal, et al.
Veröffentlicht: (2025)
von: Lancewicki, Tal, et al.
Veröffentlicht: (2025)
Learned Cost Model for Placement on Reconfigurable Dataflow Hardware
von: Guha, Etash, et al.
Veröffentlicht: (2025)
von: Guha, Etash, et al.
Veröffentlicht: (2025)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
von: Boone, Victor, et al.
Veröffentlicht: (2024)
von: Boone, Victor, et al.
Veröffentlicht: (2024)
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
von: Zamir, Guy, et al.
Veröffentlicht: (2026)
von: Zamir, Guy, et al.
Veröffentlicht: (2026)
Solving robust MDPs as a sequence of static RL problems
von: Zouitine, Adil, et al.
Veröffentlicht: (2024)
von: Zouitine, Adil, et al.
Veröffentlicht: (2024)
Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
von: Ganesh, Swetha, et al.
Veröffentlicht: (2024)
Robust Parameter Learning for Uncertain MDPs
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2026)
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2026)
Minimax Optimal Variance-Aware Regret Bounds for Multinomial Logistic MDPs
von: Boudart, Pierre, et al.
Veröffentlicht: (2026)
von: Boudart, Pierre, et al.
Veröffentlicht: (2026)
Convex Is Back: Solving Belief MDPs With Convexity-Informed Deep Reinforcement Learning
von: Koutas, Daniel, et al.
Veröffentlicht: (2025)
von: Koutas, Daniel, et al.
Veröffentlicht: (2025)
Dynamic Regret Reduces to Kernelized Static Regret
von: Jacobsen, Andrew, et al.
Veröffentlicht: (2025)
von: Jacobsen, Andrew, et al.
Veröffentlicht: (2025)
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
von: Zhang, Runyu, et al.
Veröffentlicht: (2023)
von: Zhang, Runyu, et al.
Veröffentlicht: (2023)
Policy Gradient in Robust MDPs with Global Convergence Guarantee
von: Wang, Qiuhao, et al.
Veröffentlicht: (2022)
von: Wang, Qiuhao, et al.
Veröffentlicht: (2022)
Efficient Solution and Learning of Robust Factored MDPs
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2025)
von: Schnitzer, Yannik, et al.
Veröffentlicht: (2025)
No-Regret is not enough! Bandits with General Constraints through Adaptive Regret Minimization
von: Bernasconi, Martino, et al.
Veröffentlicht: (2024)
von: Bernasconi, Martino, et al.
Veröffentlicht: (2024)
DeepAveragers: Offline Reinforcement Learning by Solving Derived Non-Parametric MDPs
von: Shrestha, Aayam, et al.
Veröffentlicht: (2020)
von: Shrestha, Aayam, et al.
Veröffentlicht: (2020)
Demystifying Linear MDPs and Novel Dynamics Aggregation Framework
von: Lee, Joongkyu, et al.
Veröffentlicht: (2024)
von: Lee, Joongkyu, et al.
Veröffentlicht: (2024)
Efficient Duple Perturbation Robustness in Low-rank MDPs
von: Hu, Yang, et al.
Veröffentlicht: (2024)
von: Hu, Yang, et al.
Veröffentlicht: (2024)
Provably Efficient Algorithms for S- and Non-Rectangular Robust MDPs with General Parameterization
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2026)
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2026)
Primal-Only Actor Critic Algorithm for Robust Constrained Average Cost MDPs
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2025)
von: Satheesh, Anirudh, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On the Diminishing Returns of Width for Continual Learning
von: Guha, Etash, et al.
Veröffentlicht: (2024) -
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
von: Li, Long-Fei, et al.
Veröffentlicht: (2024) -
Truly No-Regret Learning in Constrained MDPs
von: Müller, Adrian, et al.
Veröffentlicht: (2024) -
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
von: Gadot, Uri, et al.
Veröffentlicht: (2023) -
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
von: Wang, Kaixin, et al.
Veröffentlicht: (2023)