An Improved Model-Free Decision-Estimation Coefficient with Applications in Adversarial MDPs
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Haolin, Wei, Chen-Yu, Zimmert, Julian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
por: Liu, Haolin, et al.
Publicado: (2024)
por: Liu, Haolin, et al.
Publicado: (2024)
Decision Making in Hybrid Environments: A Model Aggregation Approach
por: Liu, Haolin, et al.
Publicado: (2025)
por: Liu, Haolin, et al.
Publicado: (2025)
A Model Selection Approach for Corruption Robust Reinforcement Learning
por: Wei, Chen-Yu, et al.
Publicado: (2021)
por: Wei, Chen-Yu, et al.
Publicado: (2021)
Optimal cross-learning for contextual bandits with unknown context distributions
por: Schneider, Jon, et al.
Publicado: (2024)
por: Schneider, Jon, et al.
Publicado: (2024)
A Best-of-both-worlds Algorithm for Bandits with Delayed Feedback with Robustness to Excessive Delays
por: Masoudian, Saeed, et al.
Publicado: (2023)
por: Masoudian, Saeed, et al.
Publicado: (2023)
Incentive-compatible Bandits: Importance Weighting No More
por: Zimmert, Julian, et al.
Publicado: (2024)
por: Zimmert, Julian, et al.
Publicado: (2024)
Unified Algorithms for RL with Decision-Estimation Coefficients: PAC, Reward-Free, Preference-Based Learning, and Beyond
por: Chen, Fan, et al.
Publicado: (2022)
por: Chen, Fan, et al.
Publicado: (2022)
Improved Algorithm for Adversarial Linear Mixture MDPs with Bandit Feedback and Unknown Transition
por: Li, Long-Fei, et al.
Publicado: (2024)
por: Li, Long-Fei, et al.
Publicado: (2024)
Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs
por: Kash, Ian A., et al.
Publicado: (2022)
por: Kash, Ian A., et al.
Publicado: (2022)
Non-stationary Bandit Convex Optimization: A Comprehensive Study
por: Liu, Xiaoqi, et al.
Publicado: (2025)
por: Liu, Xiaoqi, et al.
Publicado: (2025)
Learning Adversarial MDPs with Stochastic Hard Constraints
por: Stradi, Francesco Emanuele, et al.
Publicado: (2024)
por: Stradi, Francesco Emanuele, et al.
Publicado: (2024)
NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
por: Si, Chongjie, et al.
Publicado: (2025)
por: Si, Chongjie, et al.
Publicado: (2025)
Reinforcement Learning from Adversarial Preferences in Tabular MDPs
por: Tsuchiya, Taira, et al.
Publicado: (2025)
por: Tsuchiya, Taira, et al.
Publicado: (2025)
Online Learning in MDPs with Partially Adversarial Transitions and Losses
por: Schlisselberg, Ofir, et al.
Publicado: (2026)
por: Schlisselberg, Ofir, et al.
Publicado: (2026)
Contextual Dynamic Pricing with Heterogeneous Buyers
por: Lykouris, Thodoris, et al.
Publicado: (2025)
por: Lykouris, Thodoris, et al.
Publicado: (2025)
Efficient Model-Free Exploration in Low-Rank MDPs
por: Mhammedi, Zakaria, et al.
Publicado: (2023)
por: Mhammedi, Zakaria, et al.
Publicado: (2023)
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
por: Li, Long-Fei, et al.
Publicado: (2024)
por: Li, Long-Fei, et al.
Publicado: (2024)
Breiman meets Bellman: Non-Greedy Decision Trees with MDPs
por: Kohler, Hector, et al.
Publicado: (2023)
por: Kohler, Hector, et al.
Publicado: (2023)
$κ$-Explorer: A Unified Framework for Active Model Estimation in MDPs
por: Gu, Xihe, et al.
Publicado: (2026)
por: Gu, Xihe, et al.
Publicado: (2026)
Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
por: Ito, Shinji, et al.
Publicado: (2025)
por: Ito, Shinji, et al.
Publicado: (2025)
Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization
por: Tiapkin, Daniil, et al.
Publicado: (2024)
por: Tiapkin, Daniil, et al.
Publicado: (2024)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
por: Zhang, Junkai, et al.
Publicado: (2023)
por: Zhang, Junkai, et al.
Publicado: (2023)
TBDFiltering: Sample-Efficient Tree-Based Data Filtering
por: Busa-Fekete, Robert Istvan, et al.
Publicado: (2026)
por: Busa-Fekete, Robert Istvan, et al.
Publicado: (2026)
How Do Diffusion Models Improve Adversarial Robustness?
por: Yuezhang, Liu, et al.
Publicado: (2025)
por: Yuezhang, Liu, et al.
Publicado: (2025)
On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage
por: Liu, Haolin, et al.
Publicado: (2026)
por: Liu, Haolin, et al.
Publicado: (2026)
An Improved Algorithm for Adversarial Linear Contextual Bandits via Reduction
por: van Erven, Tim, et al.
Publicado: (2025)
por: van Erven, Tim, et al.
Publicado: (2025)
Adaptive Estimation and Optimal Control in Offline Contextual MDPs without Stationarity
por: Bhattacharyya, Riddhiman, et al.
Publicado: (2026)
por: Bhattacharyya, Riddhiman, et al.
Publicado: (2026)
Corruption-Robust Linear Bandits: Minimax Optimality and Gap-Dependent Misspecification
por: Liu, Haolin, et al.
Publicado: (2024)
por: Liu, Haolin, et al.
Publicado: (2024)
Efficient Opportunistic Approachability
por: Marinov, Teodor Vanislavov, et al.
Publicado: (2026)
por: Marinov, Teodor Vanislavov, et al.
Publicado: (2026)
Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
por: Liu, Xingtu, et al.
Publicado: (2025)
por: Liu, Xingtu, et al.
Publicado: (2025)
Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs
por: Tan, Kevin, et al.
Publicado: (2024)
por: Tan, Kevin, et al.
Publicado: (2024)
Teaching Precommitted Agents: Model-Free Policy Evaluation and Control in Quasi-Hyperbolic Discounted MDPs
por: Eshwar, S. R.
Publicado: (2025)
por: Eshwar, S. R.
Publicado: (2025)
Modes of Sequence Models and Learning Coefficients
por: Chen, Zhongtian, et al.
Publicado: (2025)
por: Chen, Zhongtian, et al.
Publicado: (2025)
Interval Estimation of Coefficients in Penalized Regression Models of Insurance Data
por: Manna, Alokesh, et al.
Publicado: (2024)
por: Manna, Alokesh, et al.
Publicado: (2024)
Time-Constrained Robust MDPs
por: Zouitine, Adil, et al.
Publicado: (2024)
por: Zouitine, Adil, et al.
Publicado: (2024)
Estimation of the Learning Coefficient Using Empirical Loss
por: Takio, Tatsuyoshi, et al.
Publicado: (2025)
por: Takio, Tatsuyoshi, et al.
Publicado: (2025)
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
por: Ding, Dongsheng, et al.
Publicado: (2023)
por: Ding, Dongsheng, et al.
Publicado: (2023)
Near-Optimal Sample Complexity for Online Constrained MDPs
por: Liu, Chang, et al.
Publicado: (2026)
por: Liu, Chang, et al.
Publicado: (2026)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
por: Hong, Kihyuk, et al.
Publicado: (2024)
por: Hong, Kihyuk, et al.
Publicado: (2024)
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
por: Wang, Kaixin, et al.
Publicado: (2023)
por: Wang, Kaixin, et al.
Publicado: (2023)
Ejemplares similares
-
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
por: Liu, Haolin, et al.
Publicado: (2024) -
Decision Making in Hybrid Environments: A Model Aggregation Approach
por: Liu, Haolin, et al.
Publicado: (2025) -
A Model Selection Approach for Corruption Robust Reinforcement Learning
por: Wei, Chen-Yu, et al.
Publicado: (2021) -
Optimal cross-learning for contextual bandits with unknown context distributions
por: Schneider, Jon, et al.
Publicado: (2024) -
A Best-of-both-worlds Algorithm for Bandits with Delayed Feedback with Robustness to Excessive Delays
por: Masoudian, Saeed, et al.
Publicado: (2023)