Transformers Provably Implement In-Context Reinforcement Learning with Policy Improvement
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Haodong, Lai, Lifeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Transformers Handle Endogeneity in In-Context Linear Regression
by: Liang, Haodong, et al.
Published: (2024)
by: Liang, Haodong, et al.
Published: (2024)
Provably Invincible Adversarial Attacks on Reinforcement Learning Systems: A Rate-Distortion Information-Theoretic Approach
by: Lu, Ziqing, et al.
Published: (2025)
by: Lu, Ziqing, et al.
Published: (2025)
Provable and Practical In-Context Policy Optimization for Self-Improvement
by: Yu, Tianrun, et al.
Published: (2026)
by: Yu, Tianrun, et al.
Published: (2026)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026)
by: Yang, Tong, et al.
Published: (2026)
Towards Monotonic Improvement in In-Context Reinforcement Learning
by: Zhang, Wenhao, et al.
Published: (2025)
by: Zhang, Wenhao, et al.
Published: (2025)
Differentially Private Two-Stage Gradient Descent for Instrumental Variable Regression
by: Liang, Haodong, et al.
Published: (2025)
by: Liang, Haodong, et al.
Published: (2025)
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
by: Anisimov, Maksim, et al.
Published: (2026)
by: Anisimov, Maksim, et al.
Published: (2026)
Blending Imitation and Reinforcement Learning for Robust Policy Improvement
by: Liu, Xuefeng, et al.
Published: (2023)
by: Liu, Xuefeng, et al.
Published: (2023)
Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning
by: Goodall, Alexander W., et al.
Published: (2025)
by: Goodall, Alexander W., et al.
Published: (2025)
Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining
by: Lin, Licong, et al.
Published: (2023)
by: Lin, Licong, et al.
Published: (2023)
Bridging Distributional and Risk-sensitive Reinforcement Learning with Provable Regret Bounds
by: Liang, Hao, et al.
Published: (2022)
by: Liang, Hao, et al.
Published: (2022)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence
by: Xiao, Minheng, et al.
Published: (2024)
by: Xiao, Minheng, et al.
Published: (2024)
Provably Efficient Exploration in Inverse Constrained Reinforcement Learning
by: Yue, Bo, et al.
Published: (2024)
by: Yue, Bo, et al.
Published: (2024)
Provable Zero-Shot Generalization in Offline Reinforcement Learning
by: Wang, Zhiyong, et al.
Published: (2025)
by: Wang, Zhiyong, et al.
Published: (2025)
Heuristic Transformer: Belief Augmented In-Context Reinforcement Learning
by: Dippel, Oliver, et al.
Published: (2025)
by: Dippel, Oliver, et al.
Published: (2025)
The Definitive Guide to Policy Gradients in Deep Reinforcement Learning: Theory, Algorithms and Implementations
by: Lehmann, Matthias
Published: (2024)
by: Lehmann, Matthias
Published: (2024)
Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
by: Kar, Avik, et al.
Published: (2024)
by: Kar, Avik, et al.
Published: (2024)
Dynamic Model Predictive Shielding for Provably Safe Reinforcement Learning
by: Banerjee, Arko, et al.
Published: (2024)
by: Banerjee, Arko, et al.
Published: (2024)
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
by: Zhang, Hongming, et al.
Published: (2023)
by: Zhang, Hongming, et al.
Published: (2023)
Random Policy Enables In-Context Reinforcement Learning within Trust Horizons
by: Chen, Weiqin, et al.
Published: (2024)
by: Chen, Weiqin, et al.
Published: (2024)
Leveraging Analytic Gradients in Provably Safe Reinforcement Learning
by: Walter, Tim, et al.
Published: (2025)
by: Walter, Tim, et al.
Published: (2025)
Neural Paging: Learning Context Management Policies for Turing-Complete Agents
by: Chen, Liang, et al.
Published: (2026)
by: Chen, Liang, et al.
Published: (2026)
Provably Efficient Action-Manipulation Attack Against Continuous Reinforcement Learning
by: Luo, Zhi, et al.
Published: (2024)
by: Luo, Zhi, et al.
Published: (2024)
In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-Thought
by: Huang, Sili, et al.
Published: (2024)
by: Huang, Sili, et al.
Published: (2024)
Towards Provable Log Density Policy Gradient
by: Katdare, Pulkit, et al.
Published: (2024)
by: Katdare, Pulkit, et al.
Published: (2024)
Learning in Context, Guided by Choice: A Reward-Free Paradigm for Reinforcement Learning with Transformers
by: Dong, Juncheng, et al.
Published: (2026)
by: Dong, Juncheng, et al.
Published: (2026)
Implicit Safe Set Algorithm for Provably Safe Reinforcement Learning
by: Zhao, Weiye, et al.
Published: (2024)
by: Zhao, Weiye, et al.
Published: (2024)
Provable Multi-Party Reinforcement Learning with Diverse Human Feedback
by: Zhong, Huiying, et al.
Published: (2024)
by: Zhong, Huiying, et al.
Published: (2024)
Provably Efficient Reward Transfer in Reinforcement Learning with Discrete Markov Decision Processes
by: Vora, Kevin, et al.
Published: (2025)
by: Vora, Kevin, et al.
Published: (2025)
Generative Pre-Trained Transformer for Symbolic Regression Base In-Context Reinforcement Learning
by: Li, Yanjie, et al.
Published: (2024)
by: Li, Yanjie, et al.
Published: (2024)
Provable Reward-Agnostic Preference-Based Reinforcement Learning
by: Zhan, Wenhao, et al.
Published: (2023)
by: Zhan, Wenhao, et al.
Published: (2023)
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
by: Bu, Dake, et al.
Published: (2025)
by: Bu, Dake, et al.
Published: (2025)
Diverse Randomized Value Functions: A Provably Pessimistic Approach for Offline Reinforcement Learning
by: Yu, Xudong, et al.
Published: (2024)
by: Yu, Xudong, et al.
Published: (2024)
Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization
by: Hao, Ruijie, et al.
Published: (2026)
by: Hao, Ruijie, et al.
Published: (2026)
HRLAIF: Improvements in Helpfulness and Harmlessness in Open-domain Reinforcement Learning From AI Feedback
by: Li, Ang, et al.
Published: (2024)
by: Li, Ang, et al.
Published: (2024)
Diffusion Models for Reinforcement Learning: Foundations, Taxonomy, and Development
by: Xu, Changfu, et al.
Published: (2025)
by: Xu, Changfu, et al.
Published: (2025)
The Formalism-Implementation Gap in Reinforcement Learning Research
by: Castro, Pablo Samuel
Published: (2025)
by: Castro, Pablo Samuel
Published: (2025)
Multi-objective Reinforcement Learning with Nonlinear Preferences: Provable Approximation for Maximizing Expected Scalarized Return
by: Peng, Nianli, et al.
Published: (2023)
by: Peng, Nianli, et al.
Published: (2023)
Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function Approximation
by: Cho, Taehyun, et al.
Published: (2024)
by: Cho, Taehyun, et al.
Published: (2024)
Similar Items
-
Transformers Handle Endogeneity in In-Context Linear Regression
by: Liang, Haodong, et al.
Published: (2024) -
Provably Invincible Adversarial Attacks on Reinforcement Learning Systems: A Rate-Distortion Information-Theoretic Approach
by: Lu, Ziqing, et al.
Published: (2025) -
Provable and Practical In-Context Policy Optimization for Self-Improvement
by: Yu, Tianrun, et al.
Published: (2026) -
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026) -
Towards Monotonic Improvement in In-Context Reinforcement Learning
by: Zhang, Wenhao, et al.
Published: (2025)