Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions
Fuente:
arXiv
Saved in:
| Main Authors: | Hait, Soumita, Li, Ping, Luo, Haipeng, Zhang, Mengxiao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Comparator-Adaptive $Φ$-Regret: Improved Bounds, Simpler Algorithms, and Applications to Games
by: Hait, Soumita, et al.
Published: (2025)
by: Hait, Soumita, et al.
Published: (2025)
Alternating Regret for Online Convex Optimization
by: Hait, Soumita, et al.
Published: (2025)
by: Hait, Soumita, et al.
Published: (2025)
The Harder Path: Last Iterate Convergence for Uncoupled Learning in Zero-Sum Games with Bandit Feedback
by: Fiegel, Côme, et al.
Published: (2026)
by: Fiegel, Côme, et al.
Published: (2026)
Efficient Contextual Bandits with Uninformed Feedback Graphs
by: Zhang, Mengxiao, et al.
Published: (2024)
by: Zhang, Mengxiao, et al.
Published: (2024)
Contextual Multinomial Logit Bandits with General Value Functions
by: Zhang, Mengxiao, et al.
Published: (2024)
by: Zhang, Mengxiao, et al.
Published: (2024)
Near-Optimal Regret in Linear MDPs with Aggregate Bandit Feedback
by: Cassel, Asaf, et al.
Published: (2024)
by: Cassel, Asaf, et al.
Published: (2024)
Contextual Linear Bandits with Delay as Payoff
by: Zhang, Mengxiao, et al.
Published: (2025)
by: Zhang, Mengxiao, et al.
Published: (2025)
Last-Iterate Convergence of Payoff-Based Independent Learning in Zero-Sum Stochastic Games
by: Chen, Zaiwei, et al.
Published: (2024)
by: Chen, Zaiwei, et al.
Published: (2024)
Last-Iterate Convergence: Zero-Sum Games and Constrained Min-Max Optimization
by: Daskalakis, Constantinos, et al.
Published: (2018)
by: Daskalakis, Constantinos, et al.
Published: (2018)
From Average-Iterate to Last-Iterate Convergence in Games: A Reduction and Its Applications
by: Cai, Yang, et al.
Published: (2025)
by: Cai, Yang, et al.
Published: (2025)
Last-Iterate Convergence Properties of Regret-Matching Algorithms in Games
by: Cai, Yang, et al.
Published: (2023)
by: Cai, Yang, et al.
Published: (2023)
Two-Player Zero-Sum Games with Bandit Feedback
by: Yılmaz, Elif, et al.
Published: (2025)
by: Yılmaz, Elif, et al.
Published: (2025)
On Separation Between Best-Iterate, Random-Iterate, and Last-Iterate Convergence of Learning in Games
by: Cai, Yang, et al.
Published: (2025)
by: Cai, Yang, et al.
Published: (2025)
Instance-Dependent Regret Bounds for Learning Two-Player Zero-Sum Games with Bandit Feedback
by: Ito, Shinji, et al.
Published: (2025)
by: Ito, Shinji, et al.
Published: (2025)
Learning Zero-Sum Linear Quadratic Games with Improved Sample Complexity and Last-Iterate Convergence
by: Wu, Jiduan, et al.
Published: (2023)
by: Wu, Jiduan, et al.
Published: (2023)
Unregularized Linear Convergence in Zero-Sum Game from Preference Feedback
by: Chen, Shulun, et al.
Published: (2025)
by: Chen, Shulun, et al.
Published: (2025)
Fast Last-Iterate Convergence of Learning in Games Requires Forgetful Algorithms
by: Cai, Yang, et al.
Published: (2024)
by: Cai, Yang, et al.
Published: (2024)
Near-Optimal Regret for Distributed Adversarial Bandits: A Black-Box Approach
by: Qiu, Hao, et al.
Published: (2026)
by: Qiu, Hao, et al.
Published: (2026)
Interaction-Grounded Learning for Contextual Markov Decision Processes with Personalized Feedback
by: Zhang, Mengxiao, et al.
Published: (2026)
by: Zhang, Mengxiao, et al.
Published: (2026)
Near-Optimal Policy Optimization for Correlated Equilibrium in General-Sum Markov Games
by: Cai, Yang, et al.
Published: (2024)
by: Cai, Yang, et al.
Published: (2024)
One Good Source is All You Need: Near-Optimal Regret for Bandits under Heterogeneous Noise
by: Bhat, Amith, et al.
Published: (2026)
by: Bhat, Amith, et al.
Published: (2026)
Nearly Minimax Optimal Submodular Maximization with Bandit Feedback
by: Tajdini, Artin, et al.
Published: (2023)
by: Tajdini, Artin, et al.
Published: (2023)
Uniform Last-Iterate Guarantee for Bandits and Reinforcement Learning
by: Liu, Junyan, et al.
Published: (2024)
by: Liu, Junyan, et al.
Published: (2024)
Last-Iterate Convergence of No-Regret Learning for Equilibria in Bargaining Games
by: Kamp, Serafina, et al.
Published: (2025)
by: Kamp, Serafina, et al.
Published: (2025)
Decentralized Online Learning in General-Sum Stackelberg Games
by: Yu, Yaolong, et al.
Published: (2024)
by: Yu, Yaolong, et al.
Published: (2024)
Provably Efficient Interactive-Grounded Learning with Personalized Reward
by: Zhang, Mengxiao, et al.
Published: (2024)
by: Zhang, Mengxiao, et al.
Published: (2024)
Adapting to Stochastic and Adversarial Losses in Episodic MDPs with Aggregate Bandit Feedback
by: Ito, Shinji, et al.
Published: (2025)
by: Ito, Shinji, et al.
Published: (2025)
Scale-Invariant Fast Convergence in Games
by: Tsuchiya, Taira, et al.
Published: (2026)
by: Tsuchiya, Taira, et al.
Published: (2026)
Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
by: Di, Qiwei, et al.
Published: (2024)
by: Di, Qiwei, et al.
Published: (2024)
Near Optimal Convergence to Coarse Correlated Equilibrium in General-Sum Markov Games
by: Yorulmaz, Asrin Efe, et al.
Published: (2025)
by: Yorulmaz, Asrin Efe, et al.
Published: (2025)
Last-Iterate Analyses of FTRL with the 1/2-Tsallis Entropy in Stochastic Bandits
by: Zhan, Jingxin, et al.
Published: (2025)
by: Zhan, Jingxin, et al.
Published: (2025)
Scale-Invariant Regret Matching and Online Learning with Optimal Convergence: Bridging Theory and Practice in Zero-Sum Games
by: Zhang, Brian Hu, et al.
Published: (2025)
by: Zhang, Brian Hu, et al.
Published: (2025)
Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback
by: Zhou, Runlong, et al.
Published: (2025)
by: Zhou, Runlong, et al.
Published: (2025)
Near-Constant Strong Violation and Last-Iterate Convergence for Online CMDPs via Decaying Safety Margins
by: Zuo, Qian, et al.
Published: (2026)
by: Zuo, Qian, et al.
Published: (2026)
On the Last-Iterate Convergence of Shuffling Gradient Methods
by: Liu, Zijian, et al.
Published: (2024)
by: Liu, Zijian, et al.
Published: (2024)
Nearly-Optimal Bandit Learning in Stackelberg Games with Side Information
by: Balcan, Maria-Florina, et al.
Published: (2025)
by: Balcan, Maria-Florina, et al.
Published: (2025)
Learning Equilibria in Matching Games with Bandit Feedback
by: Athanasopoulos, Andreas, et al.
Published: (2025)
by: Athanasopoulos, Andreas, et al.
Published: (2025)
Near-Optimal Regret in Adversarial Kernel Bandits
by: Zhang, Yu-Jie, et al.
Published: (2026)
by: Zhang, Yu-Jie, et al.
Published: (2026)
Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs
by: Lu, Michael, et al.
Published: (2026)
by: Lu, Michael, et al.
Published: (2026)
Optimal Clustering with Bandit Feedback
by: Yang, Junwen, et al.
Published: (2022)
by: Yang, Junwen, et al.
Published: (2022)
Similar Items
-
Comparator-Adaptive $Φ$-Regret: Improved Bounds, Simpler Algorithms, and Applications to Games
by: Hait, Soumita, et al.
Published: (2025) -
Alternating Regret for Online Convex Optimization
by: Hait, Soumita, et al.
Published: (2025) -
The Harder Path: Last Iterate Convergence for Uncoupled Learning in Zero-Sum Games with Bandit Feedback
by: Fiegel, Côme, et al.
Published: (2026) -
Efficient Contextual Bandits with Uninformed Feedback Graphs
by: Zhang, Mengxiao, et al.
Published: (2024) -
Contextual Multinomial Logit Bandits with General Value Functions
by: Zhang, Mengxiao, et al.
Published: (2024)