Offline-Online Reinforcement Learning for Linear Mixture MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zhongjun, Sinclair, Sean R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reinforcement Learning in MDPs with Information-Ordered Policies
by: Zhang, Zhongjun, et al.
Published: (2025)
by: Zhang, Zhongjun, et al.
Published: (2025)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
Provable Offline Reinforcement Learning for Structured Cyclic MDPs
by: Lee, Kyungbok, et al.
Published: (2026)
by: Lee, Kyungbok, et al.
Published: (2026)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
by: Zhang, Junkai, et al.
Published: (2023)
by: Zhang, Junkai, et al.
Published: (2023)
Exploiting Exogenous Structure for Sample-Efficient Reinforcement Learning
by: Wan, Jia, et al.
Published: (2024)
by: Wan, Jia, et al.
Published: (2024)
Reward-Relevance-Filtered Linear Offline Reinforcement Learning
by: Zhou, Angela
Published: (2024)
by: Zhou, Angela
Published: (2024)
Offline Reinforcement Learning via Linear-Programming with Error-Bound Induced Constraints
by: Ozdaglar, Asuman, et al.
Published: (2022)
by: Ozdaglar, Asuman, et al.
Published: (2022)
Joint MDPs and Reinforcement Learning in Coupled-Dynamics Environments
by: Kaya, Ege C., et al.
Published: (2026)
by: Kaya, Ege C., et al.
Published: (2026)
The Data-Driven Censored Newsvendor Problem
by: Hssaine, Chamsi, et al.
Published: (2024)
by: Hssaine, Chamsi, et al.
Published: (2024)
Residuals-based Offline Reinforcement Learning
by: Zhu, Qing, et al.
Published: (2026)
by: Zhu, Qing, et al.
Published: (2026)
Offline Reinforcement Learning via Inverse Optimization
by: Dimanidis, Ioannis, et al.
Published: (2025)
by: Dimanidis, Ioannis, et al.
Published: (2025)
Online Reinforcement Learning in Markov Decision Process Using Linear Programming
by: Leon, Vincent, et al.
Published: (2023)
by: Leon, Vincent, et al.
Published: (2023)
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
by: Zhang, Runyu, et al.
Published: (2023)
by: Zhang, Runyu, et al.
Published: (2023)
Non-Stationary Inventory Control with Lead Times
by: Amiri, Nele H., et al.
Published: (2026)
by: Amiri, Nele H., et al.
Published: (2026)
Online Residual Learning from Offline Experts for Pedestrian Tracking
by: Vlachos, Anastasios, et al.
Published: (2024)
by: Vlachos, Anastasios, et al.
Published: (2024)
Offline Hierarchical Reinforcement Learning via Inverse Optimization
by: Schmidt, Carolin, et al.
Published: (2024)
by: Schmidt, Carolin, et al.
Published: (2024)
Operator Models for Continuous-Time Offline Reinforcement Learning
by: Hoischen, Nicolas, et al.
Published: (2025)
by: Hoischen, Nicolas, et al.
Published: (2025)
Planning and Learning in Average Risk-aware MDPs
by: Wang, Weikai, et al.
Published: (2025)
by: Wang, Weikai, et al.
Published: (2025)
Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning
by: Di, Qiwei, et al.
Published: (2023)
by: Di, Qiwei, et al.
Published: (2023)
Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning
by: Zhang, Dake, et al.
Published: (2024)
by: Zhang, Dake, et al.
Published: (2024)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
by: Boone, Victor, et al.
Published: (2024)
by: Boone, Victor, et al.
Published: (2024)
Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
by: Zhang, Weitong, et al.
Published: (2021)
by: Zhang, Weitong, et al.
Published: (2021)
Wait-Less Offline Tuning and Re-solving for Online Decision Making
by: Sun, Jingruo, et al.
Published: (2024)
by: Sun, Jingruo, et al.
Published: (2024)
Faster Fixed-Point Methods for Multichain MDPs
by: Zurek, Matthew, et al.
Published: (2025)
by: Zurek, Matthew, et al.
Published: (2025)
Near-Optimal Primal-Dual Algorithm for Learning Linear Mixture CMDPs with Adversarial Rewards
by: Yu, Kihyun, et al.
Published: (2026)
by: Yu, Kihyun, et al.
Published: (2026)
Online Fair Allocation of Perishable Resources
by: Banerjee, Siddhartha, et al.
Published: (2024)
by: Banerjee, Siddhartha, et al.
Published: (2024)
Quantizer Design for Finite Model Approximations, Model Learning, and Quantized Q-Learning for MDPs with Unbounded Spaces
by: Bicer, Osman, et al.
Published: (2025)
by: Bicer, Osman, et al.
Published: (2025)
Towards Optimal Offline Reinforcement Learning
by: Li, Mengmeng, et al.
Published: (2025)
by: Li, Mengmeng, et al.
Published: (2025)
Efficient Model-Free Exploration in Low-Rank MDPs
by: Mhammedi, Zakaria, et al.
Published: (2023)
by: Mhammedi, Zakaria, et al.
Published: (2023)
Efficient Duple Perturbation Robustness in Low-rank MDPs
by: Hu, Yang, et al.
Published: (2024)
by: Hu, Yang, et al.
Published: (2024)
Online Linear Programming with Batching
by: Xu, Haoran, et al.
Published: (2024)
by: Xu, Haoran, et al.
Published: (2024)
Model approximation in MDPs with unbounded per-step cost
by: Bozkurt, Berk, et al.
Published: (2024)
by: Bozkurt, Berk, et al.
Published: (2024)
Learning Neural Contracting Dynamics: Extended Linearization and Global Guarantees
by: Jaffe, Sean, et al.
Published: (2024)
by: Jaffe, Sean, et al.
Published: (2024)
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
by: Ding, Dongsheng, et al.
Published: (2023)
by: Ding, Dongsheng, et al.
Published: (2023)
On the Linear Speedup of Personalized Federated Reinforcement Learning with Shared Representations
by: Xiong, Guojun, et al.
Published: (2024)
by: Xiong, Guojun, et al.
Published: (2024)
Landscape of Policy Optimization for Finite Horizon MDPs with General State and Action
by: Chen, Xin, et al.
Published: (2024)
by: Chen, Xin, et al.
Published: (2024)
Integrated Offline and Online Learning to Solve a Large Class of Scheduling Problems
by: Liu, Anbang, et al.
Published: (2025)
by: Liu, Anbang, et al.
Published: (2025)
The Sample Complexity of Online Reinforcement Learning: A Multi-model Perspective
by: Muehlebach, Michael, et al.
Published: (2025)
by: Muehlebach, Michael, et al.
Published: (2025)
Sample Complexity of the Linear Quadratic Regulator: A Reinforcement Learning Lens
by: Moghaddam, Amirreza Neshaei, et al.
Published: (2024)
by: Moghaddam, Amirreza Neshaei, et al.
Published: (2024)
Multivariate Online Linear Regression for Hierarchical Forecasting
by: Hihat, Massil, et al.
Published: (2024)
by: Hihat, Massil, et al.
Published: (2024)
Similar Items
-
Reinforcement Learning in MDPs with Information-Ordered Policies
by: Zhang, Zhongjun, et al.
Published: (2025) -
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024) -
Provable Offline Reinforcement Learning for Structured Cyclic MDPs
by: Lee, Kyungbok, et al.
Published: (2026) -
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
by: Zhang, Junkai, et al.
Published: (2023) -
Exploiting Exogenous Structure for Sample-Efficient Reinforcement Learning
by: Wan, Jia, et al.
Published: (2024)