An Improved Model-Free Decision-Estimation Coefficient with Applications in Adversarial MDPs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Haolin, Wei, Chen-Yu, Zimmert, Julian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
von: Liu, Haolin, et al.
Veröffentlicht: (2024)
von: Liu, Haolin, et al.
Veröffentlicht: (2024)
Decision Making in Hybrid Environments: A Model Aggregation Approach
von: Liu, Haolin, et al.
Veröffentlicht: (2025)
von: Liu, Haolin, et al.
Veröffentlicht: (2025)
A Model Selection Approach for Corruption Robust Reinforcement Learning
von: Wei, Chen-Yu, et al.
Veröffentlicht: (2021)
von: Wei, Chen-Yu, et al.
Veröffentlicht: (2021)
Optimal cross-learning for contextual bandits with unknown context distributions
von: Schneider, Jon, et al.
Veröffentlicht: (2024)
von: Schneider, Jon, et al.
Veröffentlicht: (2024)
A Best-of-both-worlds Algorithm for Bandits with Delayed Feedback with Robustness to Excessive Delays
von: Masoudian, Saeed, et al.
Veröffentlicht: (2023)
von: Masoudian, Saeed, et al.
Veröffentlicht: (2023)
Incentive-compatible Bandits: Importance Weighting No More
von: Zimmert, Julian, et al.
Veröffentlicht: (2024)
von: Zimmert, Julian, et al.
Veröffentlicht: (2024)
Unified Algorithms for RL with Decision-Estimation Coefficients: PAC, Reward-Free, Preference-Based Learning, and Beyond
von: Chen, Fan, et al.
Veröffentlicht: (2022)
von: Chen, Fan, et al.
Veröffentlicht: (2022)
Improved Algorithm for Adversarial Linear Mixture MDPs with Bandit Feedback and Unknown Transition
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs
von: Kash, Ian A., et al.
Veröffentlicht: (2022)
von: Kash, Ian A., et al.
Veröffentlicht: (2022)
Non-stationary Bandit Convex Optimization: A Comprehensive Study
von: Liu, Xiaoqi, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoqi, et al.
Veröffentlicht: (2025)
Learning Adversarial MDPs with Stochastic Hard Constraints
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
von: Stradi, Francesco Emanuele, et al.
Veröffentlicht: (2024)
NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
von: Si, Chongjie, et al.
Veröffentlicht: (2025)
Reinforcement Learning from Adversarial Preferences in Tabular MDPs
von: Tsuchiya, Taira, et al.
Veröffentlicht: (2025)
von: Tsuchiya, Taira, et al.
Veröffentlicht: (2025)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2026)
von: Schlisselberg, Ofir, et al.
Veröffentlicht: (2026)
Contextual Dynamic Pricing with Heterogeneous Buyers
von: Lykouris, Thodoris, et al.
Veröffentlicht: (2025)
von: Lykouris, Thodoris, et al.
Veröffentlicht: (2025)
Efficient Model-Free Exploration in Low-Rank MDPs
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2023)
von: Mhammedi, Zakaria, et al.
Veröffentlicht: (2023)
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
von: Li, Long-Fei, et al.
Veröffentlicht: (2024)
Breiman meets Bellman: Non-Greedy Decision Trees with MDPs
von: Kohler, Hector, et al.
Veröffentlicht: (2023)
von: Kohler, Hector, et al.
Veröffentlicht: (2023)
$κ$-Explorer: A Unified Framework for Active Model Estimation in MDPs
von: Gu, Xihe, et al.
Veröffentlicht: (2026)
von: Gu, Xihe, et al.
Veröffentlicht: (2026)
Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
von: Ito, Shinji, et al.
Veröffentlicht: (2025)
von: Ito, Shinji, et al.
Veröffentlicht: (2025)
Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2024)
von: Tiapkin, Daniil, et al.
Veröffentlicht: (2024)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
von: Zhang, Junkai, et al.
Veröffentlicht: (2023)
von: Zhang, Junkai, et al.
Veröffentlicht: (2023)
TBDFiltering: Sample-Efficient Tree-Based Data Filtering
von: Busa-Fekete, Robert Istvan, et al.
Veröffentlicht: (2026)
von: Busa-Fekete, Robert Istvan, et al.
Veröffentlicht: (2026)
How Do Diffusion Models Improve Adversarial Robustness?
von: Yuezhang, Liu, et al.
Veröffentlicht: (2025)
von: Yuezhang, Liu, et al.
Veröffentlicht: (2025)
On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
von: Liu, Haolin, et al.
Veröffentlicht: (2026)
An Improved Algorithm for Adversarial Linear Contextual Bandits via Reduction
von: van Erven, Tim, et al.
Veröffentlicht: (2025)
von: van Erven, Tim, et al.
Veröffentlicht: (2025)
Adaptive Estimation and Optimal Control in Offline Contextual MDPs without Stationarity
von: Bhattacharyya, Riddhiman, et al.
Veröffentlicht: (2026)
von: Bhattacharyya, Riddhiman, et al.
Veröffentlicht: (2026)
Corruption-Robust Linear Bandits: Minimax Optimality and Gap-Dependent Misspecification
von: Liu, Haolin, et al.
Veröffentlicht: (2024)
von: Liu, Haolin, et al.
Veröffentlicht: (2024)
Efficient Opportunistic Approachability
von: Marinov, Teodor Vanislavov, et al.
Veröffentlicht: (2026)
von: Marinov, Teodor Vanislavov, et al.
Veröffentlicht: (2026)
Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
von: Liu, Xingtu, et al.
Veröffentlicht: (2025)
von: Liu, Xingtu, et al.
Veröffentlicht: (2025)
Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs
von: Tan, Kevin, et al.
Veröffentlicht: (2024)
von: Tan, Kevin, et al.
Veröffentlicht: (2024)
Teaching Precommitted Agents: Model-Free Policy Evaluation and Control in Quasi-Hyperbolic Discounted MDPs
von: Eshwar, S. R.
Veröffentlicht: (2025)
von: Eshwar, S. R.
Veröffentlicht: (2025)
Modes of Sequence Models and Learning Coefficients
von: Chen, Zhongtian, et al.
Veröffentlicht: (2025)
von: Chen, Zhongtian, et al.
Veröffentlicht: (2025)
Interval Estimation of Coefficients in Penalized Regression Models of Insurance Data
von: Manna, Alokesh, et al.
Veröffentlicht: (2024)
von: Manna, Alokesh, et al.
Veröffentlicht: (2024)
Time-Constrained Robust MDPs
von: Zouitine, Adil, et al.
Veröffentlicht: (2024)
von: Zouitine, Adil, et al.
Veröffentlicht: (2024)
Estimation of the Learning Coefficient Using Empirical Loss
von: Takio, Tatsuyoshi, et al.
Veröffentlicht: (2025)
von: Takio, Tatsuyoshi, et al.
Veröffentlicht: (2025)
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
von: Ding, Dongsheng, et al.
Veröffentlicht: (2023)
von: Ding, Dongsheng, et al.
Veröffentlicht: (2023)
Near-Optimal Sample Complexity for Online Constrained MDPs
von: Liu, Chang, et al.
Veröffentlicht: (2026)
von: Liu, Chang, et al.
Veröffentlicht: (2026)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
von: Wang, Kaixin, et al.
Veröffentlicht: (2023)
von: Wang, Kaixin, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
von: Liu, Haolin, et al.
Veröffentlicht: (2024) -
Decision Making in Hybrid Environments: A Model Aggregation Approach
von: Liu, Haolin, et al.
Veröffentlicht: (2025) -
A Model Selection Approach for Corruption Robust Reinforcement Learning
von: Wei, Chen-Yu, et al.
Veröffentlicht: (2021) -
Optimal cross-learning for contextual bandits with unknown context distributions
von: Schneider, Jon, et al.
Veröffentlicht: (2024) -
A Best-of-both-worlds Algorithm for Bandits with Delayed Feedback with Robustness to Excessive Delays
von: Masoudian, Saeed, et al.
Veröffentlicht: (2023)