Sample Complexity Characterization for Linear Contextual MDPs
Fuente:
arXiv
Guardado en:
| Autores principales: | Deng, Junze, Cheng, Yuan, Zou, Shaofeng, Liang, Yingbin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
por: Jin, Ruinan, et al.
Publicado: (2026)
por: Jin, Ruinan, et al.
Publicado: (2026)
Near-Optimal Sample Complexity for Iterated CVaR Reinforcement Learning with a Generative Model
por: Deng, Zilong, et al.
Publicado: (2025)
por: Deng, Zilong, et al.
Publicado: (2025)
Breaking the Computational Barrier: Provably Efficient Actor-Critic for Low-Rank MDPs
por: Huang, Ruiquan, et al.
Publicado: (2026)
por: Huang, Ruiquan, et al.
Publicado: (2026)
Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
por: Liu, Xingtu, et al.
Publicado: (2025)
por: Liu, Xingtu, et al.
Publicado: (2025)
Unlocking the Power of Rehearsal in Continual Learning: A Theoretical Perspective
por: Deng, Junze, et al.
Publicado: (2025)
por: Deng, Junze, et al.
Publicado: (2025)
Model-Free Robust Reinforcement Learning with Sample Complexity Analysis
por: Wang, Yudan, et al.
Publicado: (2024)
por: Wang, Yudan, et al.
Publicado: (2024)
Achieving the Asymptotically Optimal Sample Complexity of Offline Reinforcement Learning: A DRO-Based Approach
por: Wang, Yue, et al.
Publicado: (2023)
por: Wang, Yue, et al.
Publicado: (2023)
Is Pure Exploitation Sufficient in Exogenous MDPs with Linear Function Approximation?
por: Liang, Hao, et al.
Publicado: (2026)
por: Liang, Hao, et al.
Publicado: (2026)
Near-Optimal Sample Complexity for Online Constrained MDPs
por: Liu, Chang, et al.
Publicado: (2026)
por: Liu, Chang, et al.
Publicado: (2026)
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
por: Zhang, Runyu, et al.
Publicado: (2023)
por: Zhang, Runyu, et al.
Publicado: (2023)
Eluder-based Regret for Stochastic Contextual MDPs
por: Levy, Orin, et al.
Publicado: (2022)
por: Levy, Orin, et al.
Publicado: (2022)
Hybrid Reinforcement Learning Breaks Sample Size Barriers in Linear MDPs
por: Tan, Kevin, et al.
Publicado: (2024)
por: Tan, Kevin, et al.
Publicado: (2024)
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
por: Wei, Yukuan, et al.
Publicado: (2025)
por: Wei, Yukuan, et al.
Publicado: (2025)
HIPO: Instruction Hierarchy via Constrained Reinforcement Learning
por: Chen, Keru, et al.
Publicado: (2026)
por: Chen, Keru, et al.
Publicado: (2026)
When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPs
por: Escamilla, Jose Efraim Aguilar, et al.
Publicado: (2026)
por: Escamilla, Jose Efraim Aguilar, et al.
Publicado: (2026)
Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions
por: Mhammedi, Zakaria
Publicado: (2024)
por: Mhammedi, Zakaria
Publicado: (2024)
Monitoring State Transitions in Markovian Systems with Sampling Cost
por: Saurav, Kumar, et al.
Publicado: (2025)
por: Saurav, Kumar, et al.
Publicado: (2025)
Span-Based Optimal Sample Complexity for Average Reward MDPs
por: Zurek, Matthew, et al.
Publicado: (2023)
por: Zurek, Matthew, et al.
Publicado: (2023)
In-Context Learning with Representations: Contextual Generalization of Trained Transformers
por: Yang, Tong, et al.
Publicado: (2024)
por: Yang, Tong, et al.
Publicado: (2024)
Adaptive Estimation and Optimal Control in Offline Contextual MDPs without Stationarity
por: Bhattacharyya, Riddhiman, et al.
Publicado: (2026)
por: Bhattacharyya, Riddhiman, et al.
Publicado: (2026)
Projection by Convolution: Optimal Sample Complexity for Reinforcement Learning in Continuous-Space MDPs
por: Maran, Davide, et al.
Publicado: (2024)
por: Maran, Davide, et al.
Publicado: (2024)
Thompson Sampling for Multi-Objective Linear Contextual Bandit
por: Park, Somangchan, et al.
Publicado: (2025)
por: Park, Somangchan, et al.
Publicado: (2025)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
por: Hong, Kihyuk, et al.
Publicado: (2024)
por: Hong, Kihyuk, et al.
Publicado: (2024)
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
por: Weltevrede, Max, et al.
Publicado: (2024)
por: Weltevrede, Max, et al.
Publicado: (2024)
Demystifying Linear MDPs and Novel Dynamics Aggregation Framework
por: Lee, Joongkyu, et al.
Publicado: (2024)
por: Lee, Joongkyu, et al.
Publicado: (2024)
Offline Oracle-Efficient Learning for Contextual MDPs via Layerwise Exploration-Exploitation Tradeoff
por: Qian, Jian, et al.
Publicado: (2024)
por: Qian, Jian, et al.
Publicado: (2024)
Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation
por: Levy, Orin, et al.
Publicado: (2026)
por: Levy, Orin, et al.
Publicado: (2026)
From Scores to Gibbs Correctors: Accelerating Uniform-Rate Discrete Diffusion Models
por: Liang, Yuchen, et al.
Publicado: (2026)
por: Liang, Yuchen, et al.
Publicado: (2026)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
por: Zurek, Matthew, et al.
Publicado: (2024)
por: Zurek, Matthew, et al.
Publicado: (2024)
Sampling Complexity of TD and PPO in RKHS
por: Zou, Lu, et al.
Publicado: (2025)
por: Zou, Lu, et al.
Publicado: (2025)
Local Linearity: the Key for No-regret Reinforcement Learning in Continuous MDPs
por: Maran, Davide, et al.
Publicado: (2024)
por: Maran, Davide, et al.
Publicado: (2024)
Near-Optimal Dynamic Regret for Adversarial Linear Mixture MDPs
por: Li, Long-Fei, et al.
Publicado: (2024)
por: Li, Long-Fei, et al.
Publicado: (2024)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
por: Cassel, Asaf, et al.
Publicado: (2024)
por: Cassel, Asaf, et al.
Publicado: (2024)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
por: Viano, Luca, et al.
Publicado: (2024)
por: Viano, Luca, et al.
Publicado: (2024)
The Sample Complexity of Multiclass and Sparse Contextual Bandits
por: Erez, Liad, et al.
Publicado: (2026)
por: Erez, Liad, et al.
Publicado: (2026)
Span-Based Optimal Sample Complexity for Weakly Communicating and General Average Reward MDPs
por: Zurek, Matthew, et al.
Publicado: (2024)
por: Zurek, Matthew, et al.
Publicado: (2024)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
por: Zhang, Zhongjun, et al.
Publicado: (2026)
por: Zhang, Zhongjun, et al.
Publicado: (2026)
The Interplay Between Interpolation and Aggregation in Regression: Optimal Sample Complexity
por: Høgsgaard, Mikael Møller, et al.
Publicado: (2026)
por: Høgsgaard, Mikael Møller, et al.
Publicado: (2026)
Robust Offline Reinforcement Learning for Non-Markovian Decision Processes
por: Huang, Ruiquan, et al.
Publicado: (2024)
por: Huang, Ruiquan, et al.
Publicado: (2024)
Non-asymptotic Convergence of Training Transformers for Next-token Prediction
por: Huang, Ruiquan, et al.
Publicado: (2024)
por: Huang, Ruiquan, et al.
Publicado: (2024)
Ejemplares similares
-
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
por: Jin, Ruinan, et al.
Publicado: (2026) -
Near-Optimal Sample Complexity for Iterated CVaR Reinforcement Learning with a Generative Model
por: Deng, Zilong, et al.
Publicado: (2025) -
Breaking the Computational Barrier: Provably Efficient Actor-Critic for Low-Rank MDPs
por: Huang, Ruiquan, et al.
Publicado: (2026) -
Sample Complexity Bounds for Linear Constrained MDPs with a Generative Model
por: Liu, Xingtu, et al.
Publicado: (2025) -
Unlocking the Power of Rehearsal in Continual Learning: A Theoretical Perspective
por: Deng, Junze, et al.
Publicado: (2025)