Saved in:
| Main Authors: | Wang, Jiuqi, Srinivasa, Jayanth, Chen, Claire, Liu, Shuze Daniel, Payani, Ali, Zhang, Shangtong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.09044 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Experience Replay Addresses Loss of Plasticity in Continual Learning
by: Wang, Jiuqi, et al.
Published: (2025)
by: Wang, Jiuqi, et al.
Published: (2025)
Doubly Optimal Policy Evaluation for Reinforcement Learning
by: Liu, Shuze Daniel, et al.
Published: (2024)
by: Liu, Shuze Daniel, et al.
Published: (2024)
Efficient Multi-Policy Evaluation for Reinforcement Learning
by: Liu, Shuze Daniel, et al.
Published: (2024)
by: Liu, Shuze Daniel, et al.
Published: (2024)
Efficient Policy Evaluation with Safety Constraint for Reinforcement Learning
by: Chen, Claire, et al.
Published: (2024)
by: Chen, Claire, et al.
Published: (2024)
Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement Learning
by: Mahadevan, Vagul, et al.
Published: (2026)
by: Mahadevan, Vagul, et al.
Published: (2026)
Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features
by: Wang, Jiuqi, et al.
Published: (2024)
by: Wang, Jiuqi, et al.
Published: (2024)
Efficient Policy Evaluation with Offline Data Informed Behavior Policy Design
by: Liu, Shuze, et al.
Published: (2023)
by: Liu, Shuze, et al.
Published: (2023)
The ODE Method for Stochastic Approximation and Reinforcement Learning with Markovian Noise
by: Liu, Shuze Daniel, et al.
Published: (2024)
by: Liu, Shuze Daniel, et al.
Published: (2024)
Towards Provable Emergence of In-Context Reinforcement Learning
by: Wang, Jiuqi, et al.
Published: (2025)
by: Wang, Jiuqi, et al.
Published: (2025)
Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes
by: Blaser, Ethan, et al.
Published: (2026)
by: Blaser, Ethan, et al.
Published: (2026)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
by: Xie, Zixuan, et al.
Published: (2026)
by: Xie, Zixuan, et al.
Published: (2026)
Transformers Can Learn Temporal Difference Methods for In-Context Reinforcement Learning
by: Wang, Jiuqi, et al.
Published: (2024)
by: Wang, Jiuqi, et al.
Published: (2024)
A Survey of In-Context Reinforcement Learning
by: Moeini, Amir, et al.
Published: (2025)
by: Moeini, Amir, et al.
Published: (2025)
MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
Towards Formalizing Reinforcement Learning Theory
by: Zhang, Shangtong
Published: (2025)
by: Zhang, Shangtong
Published: (2025)
Context Bootstrapped Reinforcement Learning
by: Agashe, Saaket, et al.
Published: (2026)
by: Agashe, Saaket, et al.
Published: (2026)
Conformal Prediction: A Theoretical Note and Benchmarking Transductive Node Classification in Graphs
by: Maneriker, Pranav, et al.
Published: (2024)
by: Maneriker, Pranav, et al.
Published: (2024)
Counterfactual Explanations for Continuous Action Reinforcement Learning
by: Dong, Shuyang, et al.
Published: (2025)
by: Dong, Shuyang, et al.
Published: (2025)
Sparse Mixture-of-Experts for Compositional Generalization: Empirical Evidence and Theoretical Foundations of Optimal Sparsity
by: Zhao, Jinze, et al.
Published: (2024)
by: Zhao, Jinze, et al.
Published: (2024)
Generalization Error Bounds for Learning under Censored Feedback
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
On the Divergence of Differential Temporal Difference Learning without Local Clocks
by: Antrobius, David, et al.
Published: (2026)
by: Antrobius, David, et al.
Published: (2026)
Revisiting a Design Choice in Gradient Temporal Difference Learning
by: Qian, Xiaochi, et al.
Published: (2023)
by: Qian, Xiaochi, et al.
Published: (2023)
Offline Two-Player Zero-Sum Markov Games with KL Regularization
by: Chen, Claire, et al.
Published: (2026)
by: Chen, Claire, et al.
Published: (2026)
Maintaining Plasticity in Deep Continual Learning
by: Dohare, Shibhansh, et al.
Published: (2023)
by: Dohare, Shibhansh, et al.
Published: (2023)
Friends in Unexpected Places: Enhancing Local Fairness in Federated Learning through Clustering
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Multi-agent Markov Entanglement
by: Chen, Shuze, et al.
Published: (2025)
by: Chen, Shuze, et al.
Published: (2025)
Extensions of Robbins-Siegmund Theorem with Applications in Reinforcement Learning
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
Almost Sure Convergence Rates of Stochastic Approximation and Reinforcement Learning via a Poisson-Moreau Drift
by: Liu, Xinyu, et al.
Published: (2026)
by: Liu, Xinyu, et al.
Published: (2026)
Convergence and Emergence of In-Context Reinforcement Learning with Chain of Thought
by: Xie, Zixuan, et al.
Published: (2026)
by: Xie, Zixuan, et al.
Published: (2026)
CRASH: Challenging Reinforcement-Learning Based Adversarial Scenarios For Safety Hardening
by: Kulkarni, Amar, et al.
Published: (2024)
by: Kulkarni, Amar, et al.
Published: (2024)
Improving the Estimation of Lifetime Effects in A/B Testing via Treatment Locality
by: Chen, Shuze, et al.
Published: (2024)
by: Chen, Shuze, et al.
Published: (2024)
Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
Finite Sample Analysis of Linear Temporal Difference Learning with Arbitrary Features
by: Xie, Zixuan, et al.
Published: (2025)
by: Xie, Zixuan, et al.
Published: (2025)
Unlocking the Power of Rehearsal in Continual Learning: A Theoretical Perspective
by: Deng, Junze, et al.
Published: (2025)
by: Deng, Junze, et al.
Published: (2025)
Almost Sure Convergence Rates and Concentration of Stochastic Approximation and Reinforcement Learning with Markovian Noise
by: Qian, Xiaochi, et al.
Published: (2024)
by: Qian, Xiaochi, et al.
Published: (2024)
Personalized Federated Fine-tuning for Heterogeneous Data: An Automatic Rank Learning Approach via Two-Level LoRA
by: Hao, Jie, et al.
Published: (2025)
by: Hao, Jie, et al.
Published: (2025)
Asymptotic and Finite Sample Analysis of Nonexpansive Stochastic Approximations with Markovian Noise
by: Blaser, Ethan, et al.
Published: (2024)
by: Blaser, Ethan, et al.
Published: (2024)
TIER: Trajectory-Invariant Execution Rewards for Multi-Step Tool Composition
by: Kulkarni, Anay, et al.
Published: (2026)
by: Kulkarni, Anay, et al.
Published: (2026)
Global Optimality and Finite Sample Analysis of Softmax Off-Policy Actor Critic under State Distribution Mismatch
by: Zhang, Shangtong, et al.
Published: (2021)
by: Zhang, Shangtong, et al.
Published: (2021)
Similar Items
-
Experience Replay Addresses Loss of Plasticity in Continual Learning
by: Wang, Jiuqi, et al.
Published: (2025) -
Doubly Optimal Policy Evaluation for Reinforcement Learning
by: Liu, Shuze Daniel, et al.
Published: (2024) -
Efficient Multi-Policy Evaluation for Reinforcement Learning
by: Liu, Shuze Daniel, et al.
Published: (2024) -
Efficient Policy Evaluation with Safety Constraint for Reinforcement Learning
by: Chen, Claire, et al.
Published: (2024) -
Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement Learning
by: Mahadevan, Vagul, et al.
Published: (2026)