Robust Offline Reinforcement Learning for Non-Markovian Decision Processes
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Ruiquan, Liang, Yingbin, Yang, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Non-asymptotic Convergence of Training Transformers for Next-token Prediction
by: Huang, Ruiquan, et al.
Published: (2024)
by: Huang, Ruiquan, et al.
Published: (2024)
Provably Efficient UCB-type Algorithms For Learning Predictive State Representations
by: Huang, Ruiquan, et al.
Published: (2023)
by: Huang, Ruiquan, et al.
Published: (2023)
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias
by: Huang, Ruiquan, et al.
Published: (2025)
by: Huang, Ruiquan, et al.
Published: (2025)
Breaking the Computational Barrier: Provably Efficient Actor-Critic for Low-Rank MDPs
by: Huang, Ruiquan, et al.
Published: (2026)
by: Huang, Ruiquan, et al.
Published: (2026)
Flow Matching for Offline Reinforcement Learning with Discrete Actions
by: Khan, Fairoz Nower, et al.
Published: (2026)
by: Khan, Fairoz Nower, et al.
Published: (2026)
Augmenting Online RL with Offline Data is All You Need: A Unified Hybrid RL Algorithm Design and Analysis
by: Huang, Ruiquan, et al.
Published: (2025)
by: Huang, Ruiquan, et al.
Published: (2025)
Model-Based Reinforcement Learning in Discrete-Action Non-Markovian Reward Decision Processes
by: Trapasso, Alessandro, et al.
Published: (2025)
by: Trapasso, Alessandro, et al.
Published: (2025)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026)
by: Yang, Tong, et al.
Published: (2026)
Monitoring State Transitions in Markovian Systems with Sampling Cost
by: Saurav, Kumar, et al.
Published: (2025)
by: Saurav, Kumar, et al.
Published: (2025)
Solving Continual Offline Reinforcement Learning with Decision Transformer
by: Huang, Kaixin, et al.
Published: (2024)
by: Huang, Kaixin, et al.
Published: (2024)
Federated Online Prediction from Experts with Differential Privacy: Separations and Regret Speed-ups
by: Gao, Fengyu, et al.
Published: (2024)
by: Gao, Fengyu, et al.
Published: (2024)
Epistemic Robust Offline Reinforcement Learning
by: Chenreddy, Abhilash Reddy, et al.
Published: (2026)
by: Chenreddy, Abhilash Reddy, et al.
Published: (2026)
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation
by: Pan, Pei-Chi, et al.
Published: (2026)
by: Pan, Pei-Chi, et al.
Published: (2026)
A Non-Monolithic Policy Approach of Offline-to-Online Reinforcement Learning
by: Kim, JaeYoon, et al.
Published: (2024)
by: Kim, JaeYoon, et al.
Published: (2024)
Reinforcement Learning in Non-Markovian Environments
by: Chandak, Siddharth, et al.
Published: (2022)
by: Chandak, Siddharth, et al.
Published: (2022)
Corruption-Robust Offline Reinforcement Learning with General Function Approximation
by: Ye, Chenlu, et al.
Published: (2023)
by: Ye, Chenlu, et al.
Published: (2023)
Sparse Offline Reinforcement Learning with Corruption Robustness
by: Tran, Nam Phuong, et al.
Published: (2025)
by: Tran, Nam Phuong, et al.
Published: (2025)
Offline Trajectory Optimization for Offline Reinforcement Learning
by: Zhao, Ziqi, et al.
Published: (2024)
by: Zhao, Ziqi, et al.
Published: (2024)
Q-value Regularized Decision ConvFormer for Offline Reinforcement Learning
by: Yan, Teng, et al.
Published: (2024)
by: Yan, Teng, et al.
Published: (2024)
Bayesian Inverse Reinforcement Learning for Non-Markovian Rewards
by: Topper, Noah, et al.
Published: (2024)
by: Topper, Noah, et al.
Published: (2024)
Solving Offline Reinforcement Learning with Decision Tree Regression
by: Koirala, Prajwal, et al.
Published: (2024)
by: Koirala, Prajwal, et al.
Published: (2024)
Is Mamba Compatible with Trajectory Optimization in Offline Reinforcement Learning?
by: Dai, Yang, et al.
Published: (2024)
by: Dai, Yang, et al.
Published: (2024)
Sample Complexity of Offline Distributionally Robust Linear Markov Decision Processes
by: Wang, He, et al.
Published: (2024)
by: Wang, He, et al.
Published: (2024)
Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization
by: Zhan, Simon Sinong, et al.
Published: (2025)
by: Zhan, Simon Sinong, et al.
Published: (2025)
Near-Optimal Partially Observable Reinforcement Learning with Partial Online State Information
by: Shi, Ming, et al.
Published: (2023)
by: Shi, Ming, et al.
Published: (2023)
Regret Bounds for Reinforcement Learning from Multi-Source Imperfect Preferences
by: Shi, Ming, et al.
Published: (2026)
by: Shi, Ming, et al.
Published: (2026)
Towards Robust Offline-to-Online Reinforcement Learning via Uncertainty and Smoothness
by: Wen, Xiaoyu, et al.
Published: (2023)
by: Wen, Xiaoyu, et al.
Published: (2023)
Beyond Simple Sum of Delayed Rewards: Non-Markovian Reward Modeling for Reinforcement Learning
by: Tang, Yuting, et al.
Published: (2024)
by: Tang, Yuting, et al.
Published: (2024)
Dual-Robust Cross-Domain Offline Reinforcement Learning Against Dynamics Shifts
by: Qiao, Zhongjian, et al.
Published: (2025)
by: Qiao, Zhongjian, et al.
Published: (2025)
Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes
by: Lu, Miao, et al.
Published: (2022)
by: Lu, Miao, et al.
Published: (2022)
Policy Gradient Methods for Non-Markovian Reinforcement Learning
by: Kar, Avik, et al.
Published: (2026)
by: Kar, Avik, et al.
Published: (2026)
HarmoDT: Harmony Multi-Task Decision Transformer for Offline Reinforcement Learning
by: Hu, Shengchao, et al.
Published: (2024)
by: Hu, Shengchao, et al.
Published: (2024)
Deep Transfer $Q$-Learning for Offline Non-Stationary Reinforcement Learning
by: Chai, Jinhang, et al.
Published: (2025)
by: Chai, Jinhang, et al.
Published: (2025)
Safe Reinforcement Learning with Learned Non-Markovian Safety Constraints
by: Low, Siow Meng, et al.
Published: (2024)
by: Low, Siow Meng, et al.
Published: (2024)
Tractable Offline Learning of Regular Decision Processes
by: Deb, Ahana, et al.
Published: (2024)
by: Deb, Ahana, et al.
Published: (2024)
Constrained Policy Optimization with Explicit Behavior Density for Offline Reinforcement Learning
by: Zhang, Jing, et al.
Published: (2023)
by: Zhang, Jing, et al.
Published: (2023)
Corruption Robust Offline Reinforcement Learning with Human Feedback
by: Mandal, Debmalya, et al.
Published: (2024)
by: Mandal, Debmalya, et al.
Published: (2024)
Robust Bandwidth Estimation for Real-Time Communication with Offline Reinforcement Learning
by: Kai, Jian, et al.
Published: (2025)
by: Kai, Jian, et al.
Published: (2025)
Robust Probabilistic Shielding for Safe Offline Reinforcement Learning
by: Galesloot, Maris F. L., et al.
Published: (2026)
by: Galesloot, Maris F. L., et al.
Published: (2026)
Session-Level Dynamic Ad Load Optimization using Offline Robust Reinforcement Learning
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
Similar Items
-
Non-asymptotic Convergence of Training Transformers for Next-token Prediction
by: Huang, Ruiquan, et al.
Published: (2024) -
Provably Efficient UCB-type Algorithms For Learning Predictive State Representations
by: Huang, Ruiquan, et al.
Published: (2023) -
How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias
by: Huang, Ruiquan, et al.
Published: (2025) -
Breaking the Computational Barrier: Provably Efficient Actor-Critic for Low-Rank MDPs
by: Huang, Ruiquan, et al.
Published: (2026) -
Flow Matching for Offline Reinforcement Learning with Discrete Actions
by: Khan, Fairoz Nower, et al.
Published: (2026)