An Improved Model-Free Decision-Estimation Coefficient with Applications in Adversarial MDPs
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Haolin, Wei, Chen-Yu, Zimmert, Julian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
di: Liu, Haolin, et al.
Pubblicazione: (2024)
di: Liu, Haolin, et al.
Pubblicazione: (2024)
Decision Making in Hybrid Environments: A Model Aggregation Approach
di: Liu, Haolin, et al.
Pubblicazione: (2025)
di: Liu, Haolin, et al.
Pubblicazione: (2025)
A Model Selection Approach for Corruption Robust Reinforcement Learning
di: Wei, Chen-Yu, et al.
Pubblicazione: (2021)
di: Wei, Chen-Yu, et al.
Pubblicazione: (2021)
Optimal cross-learning for contextual bandits with unknown context distributions
di: Schneider, Jon, et al.
Pubblicazione: (2024)
di: Schneider, Jon, et al.
Pubblicazione: (2024)
A Best-of-both-worlds Algorithm for Bandits with Delayed Feedback with Robustness to Excessive Delays
di: Masoudian, Saeed, et al.
Pubblicazione: (2023)
di: Masoudian, Saeed, et al.
Pubblicazione: (2023)
Incentive-compatible Bandits: Importance Weighting No More
di: Zimmert, Julian, et al.
Pubblicazione: (2024)
di: Zimmert, Julian, et al.
Pubblicazione: (2024)
Unified Algorithms for RL with Decision-Estimation Coefficients: PAC, Reward-Free, Preference-Based Learning, and Beyond
di: Chen, Fan, et al.
Pubblicazione: (2022)
di: Chen, Fan, et al.
Pubblicazione: (2022)
Improved Algorithm for Adversarial Linear Mixture MDPs with Bandit Feedback and Unknown Transition
di: Li, Long-Fei, et al.
Pubblicazione: (2024)
di: Li, Long-Fei, et al.
Pubblicazione: (2024)
Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs
di: Kash, Ian A., et al.
Pubblicazione: (2022)
di: Kash, Ian A., et al.
Pubblicazione: (2022)
Non-stationary Bandit Convex Optimization: A Comprehensive Study
di: Liu, Xiaoqi, et al.
Pubblicazione: (2025)
di: Liu, Xiaoqi, et al.
Pubblicazione: (2025)
Learning Adversarial MDPs with Stochastic Hard Constraints
di: Stradi, Francesco Emanuele, et al.
Pubblicazione: (2024)
di: Stradi, Francesco Emanuele, et al.
Pubblicazione: (2024)
NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
di: Si, Chongjie, et al.
Pubblicazione: (2025)
di: Si, Chongjie, et al.
Pubblicazione: (2025)
Reinforcement Learning from Adversarial Preferences in Tabular MDPs
di: Tsuchiya, Taira, et al.
Pubblicazione: (2025)
di: Tsuchiya, Taira, et al.
Pubblicazione: (2025)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
di: Schlisselberg, Ofir, et al.
Pubblicazione: (2026)
di: Schlisselberg, Ofir, et al.
Pubblicazione: (2026)
Contextual Dynamic Pricing with Heterogeneous Buyers
di: Lykouris, Thodoris, et al.
Pubblicazione: (2025)
di: Lykouris, Thodoris, et al.
Pubblicazione: (2025)
Efficient Model-Free Exploration in Low-Rank MDPs
di: Mhammedi, Zakaria, et al.
Pubblicazione: (2023)
di: Mhammedi, Zakaria, et al.
Pubblicazione: (2023)
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
di: Li, Long-Fei, et al.
Pubblicazione: (2024)
di: Li, Long-Fei, et al.
Pubblicazione: (2024)
Breiman meets Bellman: Non-Greedy Decision Trees with MDPs
di: Kohler, Hector, et al.
Pubblicazione: (2023)
di: Kohler, Hector, et al.
Pubblicazione: (2023)
$κ$-Explorer: A Unified Framework for Active Model Estimation in MDPs
di: Gu, Xihe, et al.
Pubblicazione: (2026)
di: Gu, Xihe, et al.
Pubblicazione: (2026)
Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
di: Ito, Shinji, et al.
Pubblicazione: (2025)
di: Ito, Shinji, et al.
Pubblicazione: (2025)
Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization
di: Tiapkin, Daniil, et al.
Pubblicazione: (2024)
di: Tiapkin, Daniil, et al.
Pubblicazione: (2024)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
di: Zhang, Junkai, et al.
Pubblicazione: (2023)
di: Zhang, Junkai, et al.
Pubblicazione: (2023)
TBDFiltering: Sample-Efficient Tree-Based Data Filtering
di: Busa-Fekete, Robert Istvan, et al.
Pubblicazione: (2026)
di: Busa-Fekete, Robert Istvan, et al.
Pubblicazione: (2026)
How Do Diffusion Models Improve Adversarial Robustness?
di: Yuezhang, Liu, et al.
Pubblicazione: (2025)
di: Yuezhang, Liu, et al.
Pubblicazione: (2025)
On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage
di: Liu, Haolin, et al.
Pubblicazione: (2026)
di: Liu, Haolin, et al.
Pubblicazione: (2026)
An Improved Algorithm for Adversarial Linear Contextual Bandits via Reduction
di: van Erven, Tim, et al.
Pubblicazione: (2025)
di: van Erven, Tim, et al.
Pubblicazione: (2025)
Adaptive Estimation and Optimal Control in Offline Contextual MDPs without Stationarity
di: Bhattacharyya, Riddhiman, et al.
Pubblicazione: (2026)
di: Bhattacharyya, Riddhiman, et al.
Pubblicazione: (2026)
Corruption-Robust Linear Bandits: Minimax Optimality and Gap-Dependent Misspecification
di: Liu, Haolin, et al.
Pubblicazione: (2024)
di: Liu, Haolin, et al.
Pubblicazione: (2024)
Efficient Opportunistic Approachability
di: Marinov, Teodor Vanislavov, et al.
Pubblicazione: (2026)
di: Marinov, Teodor Vanislavov, et al.
Pubblicazione: (2026)
Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
di: Liu, Xingtu, et al.
Pubblicazione: (2025)
di: Liu, Xingtu, et al.
Pubblicazione: (2025)
Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs
di: Tan, Kevin, et al.
Pubblicazione: (2024)
di: Tan, Kevin, et al.
Pubblicazione: (2024)
Teaching Precommitted Agents: Model-Free Policy Evaluation and Control in Quasi-Hyperbolic Discounted MDPs
di: Eshwar, S. R.
Pubblicazione: (2025)
di: Eshwar, S. R.
Pubblicazione: (2025)
Modes of Sequence Models and Learning Coefficients
di: Chen, Zhongtian, et al.
Pubblicazione: (2025)
di: Chen, Zhongtian, et al.
Pubblicazione: (2025)
Interval Estimation of Coefficients in Penalized Regression Models of Insurance Data
di: Manna, Alokesh, et al.
Pubblicazione: (2024)
di: Manna, Alokesh, et al.
Pubblicazione: (2024)
Time-Constrained Robust MDPs
di: Zouitine, Adil, et al.
Pubblicazione: (2024)
di: Zouitine, Adil, et al.
Pubblicazione: (2024)
Estimation of the Learning Coefficient Using Empirical Loss
di: Takio, Tatsuyoshi, et al.
Pubblicazione: (2025)
di: Takio, Tatsuyoshi, et al.
Pubblicazione: (2025)
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
di: Ding, Dongsheng, et al.
Pubblicazione: (2023)
di: Ding, Dongsheng, et al.
Pubblicazione: (2023)
Near-Optimal Sample Complexity for Online Constrained MDPs
di: Liu, Chang, et al.
Pubblicazione: (2026)
di: Liu, Chang, et al.
Pubblicazione: (2026)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
di: Hong, Kihyuk, et al.
Pubblicazione: (2024)
di: Hong, Kihyuk, et al.
Pubblicazione: (2024)
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
di: Wang, Kaixin, et al.
Pubblicazione: (2023)
di: Wang, Kaixin, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
di: Liu, Haolin, et al.
Pubblicazione: (2024) -
Decision Making in Hybrid Environments: A Model Aggregation Approach
di: Liu, Haolin, et al.
Pubblicazione: (2025) -
A Model Selection Approach for Corruption Robust Reinforcement Learning
di: Wei, Chen-Yu, et al.
Pubblicazione: (2021) -
Optimal cross-learning for contextual bandits with unknown context distributions
di: Schneider, Jon, et al.
Pubblicazione: (2024) -
A Best-of-both-worlds Algorithm for Bandits with Delayed Feedback with Robustness to Excessive Delays
di: Masoudian, Saeed, et al.
Pubblicazione: (2023)