Uniform Last-Iterate Guarantee for Bandits and Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Junyan, Li, Yunfan, Wang, Ruosong, Yang, Lin F. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Model-Misspecification in Reinforcement Learning
von: Li, Yunfan, et al.
Veröffentlicht: (2023)
von: Li, Yunfan, et al.
Veröffentlicht: (2023)
List Replicable Reinforcement Learning
von: Zhang, Bohan, et al.
Veröffentlicht: (2025)
von: Zhang, Bohan, et al.
Veröffentlicht: (2025)
Frozen Policy Iteration: Computationally Efficient RL under Linear $Q^π$ Realizability for Deterministic Dynamics
von: Ke, Yijing, et al.
Veröffentlicht: (2026)
von: Ke, Yijing, et al.
Veröffentlicht: (2026)
Last-Iterate Guarantees for Learning in Co-coercive Games
von: Chandak, Siddharth, et al.
Veröffentlicht: (2026)
von: Chandak, Siddharth, et al.
Veröffentlicht: (2026)
Principal-Agent Bandit Games with Self-Interested and Exploratory Learning Agents
von: Liu, Junyan, et al.
Veröffentlicht: (2024)
von: Liu, Junyan, et al.
Veröffentlicht: (2024)
The Harder Path: Last Iterate Convergence for Uncoupled Learning in Zero-Sum Games with Bandit Feedback
von: Fiegel, Côme, et al.
Veröffentlicht: (2026)
von: Fiegel, Côme, et al.
Veröffentlicht: (2026)
Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning
von: Wang, Yiran, et al.
Veröffentlicht: (2024)
von: Wang, Yiran, et al.
Veröffentlicht: (2024)
Misspecified $Q$-Learning with Sparse Linear Function Approximation: Tight Bounds on Approximation Error
von: Du, Ally Yalei, et al.
Veröffentlicht: (2024)
von: Du, Ally Yalei, et al.
Veröffentlicht: (2024)
The Bandit Whisperer: Communication Learning for Restless Bandits
von: Zhao, Yunfan, et al.
Veröffentlicht: (2024)
von: Zhao, Yunfan, et al.
Veröffentlicht: (2024)
Last-Iterate Analyses of FTRL with the 1/2-Tsallis Entropy in Stochastic Bandits
von: Zhan, Jingxin, et al.
Veröffentlicht: (2025)
von: Zhan, Jingxin, et al.
Veröffentlicht: (2025)
Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions
von: Hait, Soumita, et al.
Veröffentlicht: (2026)
von: Hait, Soumita, et al.
Veröffentlicht: (2026)
Last-Iterate Global Convergence of Policy Gradients for Constrained Reinforcement Learning
von: Montenegro, Alessandro, et al.
Veröffentlicht: (2024)
von: Montenegro, Alessandro, et al.
Veröffentlicht: (2024)
BanditQ: Fair Bandits with Guaranteed Rewards
von: Sinha, Abhishek
Veröffentlicht: (2023)
von: Sinha, Abhishek
Veröffentlicht: (2023)
Near-Optimal Sample Complexity for Online Constrained MDPs
von: Liu, Chang, et al.
Veröffentlicht: (2026)
von: Liu, Chang, et al.
Veröffentlicht: (2026)
Learning for Bandits under Action Erasures
von: Hanna, Osama, et al.
Veröffentlicht: (2024)
von: Hanna, Osama, et al.
Veröffentlicht: (2024)
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs
von: Lu, Michael, et al.
Veröffentlicht: (2026)
von: Lu, Michael, et al.
Veröffentlicht: (2026)
Fairness and Privacy Guarantees in Federated Contextual Bandits
von: Solanki, Sambhav, et al.
Veröffentlicht: (2024)
von: Solanki, Sambhav, et al.
Veröffentlicht: (2024)
On the Last-Iterate Convergence of Shuffling Gradient Methods
von: Liu, Zijian, et al.
Veröffentlicht: (2024)
von: Liu, Zijian, et al.
Veröffentlicht: (2024)
Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee
von: Hung, Yu-Heng, et al.
Veröffentlicht: (2025)
von: Hung, Yu-Heng, et al.
Veröffentlicht: (2025)
Pure Exploration for a Good Policy in Reinforcement Learning with Bandit Feedback
von: Li, Zitian, et al.
Veröffentlicht: (2026)
von: Li, Zitian, et al.
Veröffentlicht: (2026)
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
On Separation Between Best-Iterate, Random-Iterate, and Last-Iterate Convergence of Learning in Games
von: Cai, Yang, et al.
Veröffentlicht: (2025)
von: Cai, Yang, et al.
Veröffentlicht: (2025)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
von: Lu, Xiaodong, et al.
Veröffentlicht: (2026)
von: Lu, Xiaodong, et al.
Veröffentlicht: (2026)
Dynamic Learning Rate for Deep Reinforcement Learning: A Bandit Approach
von: Donâncio, Henrique, et al.
Veröffentlicht: (2024)
von: Donâncio, Henrique, et al.
Veröffentlicht: (2024)
Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond
von: Liu, Xutong, et al.
Veröffentlicht: (2024)
von: Liu, Xutong, et al.
Veröffentlicht: (2024)
Revisiting the Last-Iterate Convergence of Stochastic Gradient Methods
von: Liu, Zijian, et al.
Veröffentlicht: (2023)
von: Liu, Zijian, et al.
Veröffentlicht: (2023)
Attack-Resistant Uniform Fairness for Linear and Smooth Contextual Bandits
von: Zhang, Qingwen, et al.
Veröffentlicht: (2026)
von: Zhang, Qingwen, et al.
Veröffentlicht: (2026)
Feasible Policy Iteration for Safe Reinforcement Learning
von: Yang, Yujie, et al.
Veröffentlicht: (2023)
von: Yang, Yujie, et al.
Veröffentlicht: (2023)
Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning
von: Liao, Luofeng, et al.
Veröffentlicht: (2021)
von: Liao, Luofeng, et al.
Veröffentlicht: (2021)
Last Iterate Convergence of Incremental Methods and Applications in Continual Learning
von: Cai, Xufeng, et al.
Veröffentlicht: (2024)
von: Cai, Xufeng, et al.
Veröffentlicht: (2024)
Sharper Guarantees for Misspecified Kernelized Bandit Optimization
von: Maran, Davide, et al.
Veröffentlicht: (2026)
von: Maran, Davide, et al.
Veröffentlicht: (2026)
Pairwise Elimination with Instance-Dependent Guarantees for Bandits with Cost Subsidy
von: Juneja, Ishank, et al.
Veröffentlicht: (2025)
von: Juneja, Ishank, et al.
Veröffentlicht: (2025)
Networked Restless Multi-Arm Bandits with Reinforcement Learning
von: Zhang, Hanmo, et al.
Veröffentlicht: (2025)
von: Zhang, Hanmo, et al.
Veröffentlicht: (2025)
Geometry-Aware Approaches for Balancing Performance and Theoretical Guarantees in Linear Bandits
von: Luo, Yuwei, et al.
Veröffentlicht: (2023)
von: Luo, Yuwei, et al.
Veröffentlicht: (2023)
Algorithm Design and Stronger Guarantees for the Improving Multi-Armed Bandits Problem
von: Blum, Avrim, et al.
Veröffentlicht: (2025)
von: Blum, Avrim, et al.
Veröffentlicht: (2025)
Provably Efficient Reinforcement Learning for Adversarial Restless Multi-Armed Bandits with Unknown Transitions and Bandit Feedback
von: Xiong, Guojun, et al.
Veröffentlicht: (2024)
von: Xiong, Guojun, et al.
Veröffentlicht: (2024)
Exponential Convergence Guarantees for Iterative Markovian Fitting
von: Silveri, Marta Gentiloni, et al.
Veröffentlicht: (2025)
von: Silveri, Marta Gentiloni, et al.
Veröffentlicht: (2025)
On the Role of Iterative Computation in Reinforcement Learning
von: Ghugare, Raj, et al.
Veröffentlicht: (2026)
von: Ghugare, Raj, et al.
Veröffentlicht: (2026)
Last-Iterate Convergence of No-Regret Learning for Equilibria in Bargaining Games
von: Kamp, Serafina, et al.
Veröffentlicht: (2025)
von: Kamp, Serafina, et al.
Veröffentlicht: (2025)
Pessimistic Value Iteration for Multi-Task Data Sharing in Offline Reinforcement Learning
von: Bai, Chenjia, et al.
Veröffentlicht: (2024)
von: Bai, Chenjia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On the Model-Misspecification in Reinforcement Learning
von: Li, Yunfan, et al.
Veröffentlicht: (2023) -
List Replicable Reinforcement Learning
von: Zhang, Bohan, et al.
Veröffentlicht: (2025) -
Frozen Policy Iteration: Computationally Efficient RL under Linear $Q^π$ Realizability for Deterministic Dynamics
von: Ke, Yijing, et al.
Veröffentlicht: (2026) -
Last-Iterate Guarantees for Learning in Co-coercive Games
von: Chandak, Siddharth, et al.
Veröffentlicht: (2026) -
Principal-Agent Bandit Games with Self-Interested and Exploratory Learning Agents
von: Liu, Junyan, et al.
Veröffentlicht: (2024)