Provable Offline Reinforcement Learning for Structured Cyclic MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Kyungbok, Sarteau, Angelica Cristello, Kosorok, Michael R. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026)
by: Yang, Tong, et al.
Published: (2026)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
by: Zhang, Zhongjun, et al.
Published: (2026)
by: Zhang, Zhongjun, et al.
Published: (2026)
Seeing Through Risk: A Symbolic Approximation of Prospect Theory
by: Yousaf, Ali Arslan, et al.
Published: (2025)
by: Yousaf, Ali Arslan, et al.
Published: (2025)
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence
by: Xiao, Minheng, et al.
Published: (2024)
by: Xiao, Minheng, et al.
Published: (2024)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
by: Ding, Dongsheng, et al.
Published: (2022)
by: Ding, Dongsheng, et al.
Published: (2022)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
Structured Difference-of-Q via Orthogonal Learning
by: Cao, Defu, et al.
Published: (2024)
by: Cao, Defu, et al.
Published: (2024)
Recursive Entropic Risk Optimization in Discounted MDPs: Sample Complexity Bounds with a Generative Model
by: Mortensen, Oliver, et al.
Published: (2025)
by: Mortensen, Oliver, et al.
Published: (2025)
DCILP: A Distributed Approach for Large-Scale Causal Structure Learning
by: Dong, Shuyu, et al.
Published: (2024)
by: Dong, Shuyu, et al.
Published: (2024)
Integrated Offline and Online Learning to Solve a Large Class of Scheduling Problems
by: Liu, Anbang, et al.
Published: (2025)
by: Liu, Anbang, et al.
Published: (2025)
Provable Acceleration for Diffusion Models under Minimal Assumptions
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Learning Large Causal Structures from Inverse Covariance Matrix via Sparse Matrix Decomposition
by: Dong, Shuyu, et al.
Published: (2022)
by: Dong, Shuyu, et al.
Published: (2022)
Locally Adaptive Multi-Objective Learning
by: Kaur, Jivat Neet, et al.
Published: (2026)
by: Kaur, Jivat Neet, et al.
Published: (2026)
Combining Reinforcement Learning and Optimal Transport for the Traveling Salesman Problem
by: Goh, Yong Liang, et al.
Published: (2022)
by: Goh, Yong Liang, et al.
Published: (2022)
Provably Safe Generative Sampling with Constricting Barrier Functions
by: Gadginmath, Darshan, et al.
Published: (2026)
by: Gadginmath, Darshan, et al.
Published: (2026)
SMiLE: Provably Enforcing Global Relational Properties in Neural Networks
by: Francobaldi, Matteo, et al.
Published: (2025)
by: Francobaldi, Matteo, et al.
Published: (2025)
Faster Reinforcement Learning by Freezing Slow States
by: Wang, Yijia, et al.
Published: (2023)
by: Wang, Yijia, et al.
Published: (2023)
Accelerating Cutting-Plane Algorithms via Reinforcement Learning Surrogates
by: Mana, Kyle, et al.
Published: (2023)
by: Mana, Kyle, et al.
Published: (2023)
Graph Neural Networks for the Offline Nanosatellite Task Scheduling Problem
by: Pacheco, Bruno Machado, et al.
Published: (2023)
by: Pacheco, Bruno Machado, et al.
Published: (2023)
R2L: Reliable Reinforcement Learning: Guaranteed Return & Reliable Policies in Reinforcement Learning
by: Farhi, Nadir
Published: (2025)
by: Farhi, Nadir
Published: (2025)
Solving Truly Massive Budgeted Monotonic POMDPs with Oracle-Guided Meta-Reinforcement Learning
by: Vora, Manav, et al.
Published: (2024)
by: Vora, Manav, et al.
Published: (2024)
Infinite-Horizon Reach-Avoid Zero-Sum Games via Deep Reinforcement Learning
by: Li, Jingqi, et al.
Published: (2022)
by: Li, Jingqi, et al.
Published: (2022)
When and How Unlabeled Data Provably Improve In-Context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes
by: Lu, Miao, et al.
Published: (2022)
by: Lu, Miao, et al.
Published: (2022)
Contextual Bilevel Reinforcement Learning for Incentive Alignment
by: Thoma, Vinzenz, et al.
Published: (2024)
by: Thoma, Vinzenz, et al.
Published: (2024)
Deep Reinforcement Learning for Traveling Purchaser Problems
by: Yuan, Haofeng, et al.
Published: (2024)
by: Yuan, Haofeng, et al.
Published: (2024)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
Correlated Noise Provably Beats Independent Noise for Differentially Private Learning
by: Choquette-Choo, Christopher A., et al.
Published: (2023)
by: Choquette-Choo, Christopher A., et al.
Published: (2023)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
by: Liu, Xin, et al.
Published: (2022)
by: Liu, Xin, et al.
Published: (2022)
Adaptive Primal-Dual Method for Safe Reinforcement Learning
by: Chen, Weiqin, et al.
Published: (2024)
by: Chen, Weiqin, et al.
Published: (2024)
Performative Policy Gradient: Optimality in Performative Reinforcement Learning
by: Basu, Debabrota, et al.
Published: (2025)
by: Basu, Debabrota, et al.
Published: (2025)
Reinforcement Learning for Multi-Truck Vehicle Routing Problems
by: Levin, Joshua, et al.
Published: (2022)
by: Levin, Joshua, et al.
Published: (2022)
Hyperparameter Optimization for Driving Strategies Based on Reinforcement Learning
by: Adde, Nihal Acharya, et al.
Published: (2024)
by: Adde, Nihal Acharya, et al.
Published: (2024)
Random Pareto front surfaces
by: Tu, Ben, et al.
Published: (2024)
by: Tu, Ben, et al.
Published: (2024)
Statistical Performance Guarantee for Subgroup Identification with Generic Machine Learning
by: Li, Michael Lingzhi, et al.
Published: (2023)
by: Li, Michael Lingzhi, et al.
Published: (2023)
Provable Multi-Party Reinforcement Learning with Diverse Human Feedback
by: Zhong, Huiying, et al.
Published: (2024)
by: Zhong, Huiying, et al.
Published: (2024)
Global Group Fairness in Federated Learning via Function Tracking
by: Rychener, Yves, et al.
Published: (2025)
by: Rychener, Yves, et al.
Published: (2025)
Deterministic Policy Gradient for Reinforcement Learning with Continuous Time and State
by: Cheng, Ziheng, et al.
Published: (2025)
by: Cheng, Ziheng, et al.
Published: (2025)
Robust Learning Rate Selection for Stochastic Optimization via Splitting Diagnostic
by: Sordello, Matteo, et al.
Published: (2019)
by: Sordello, Matteo, et al.
Published: (2019)
Similar Items
-
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026) -
Offline-Online Reinforcement Learning for Linear Mixture MDPs
by: Zhang, Zhongjun, et al.
Published: (2026) -
Seeing Through Risk: A Symbolic Approximation of Prospect Theory
by: Yousaf, Ali Arslan, et al.
Published: (2025) -
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence
by: Xiao, Minheng, et al.
Published: (2024) -
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)