The Harder Path: Last Iterate Convergence for Uncoupled Learning in Zero-Sum Games with Bandit Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Fiegel, Côme, Ménard, Pierre, Kozuno, Tadashi, Valko, Michal, Perchet, Vianney |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimal last-iterate convergence in matrix games with bandit feedback using the log-barrier
by: Fiegel, Come, et al.
Published: (2026)
by: Fiegel, Come, et al.
Published: (2026)
Learning to Allocate Resources with Censored Feedback
by: Montanari, Giovanni, et al.
Published: (2026)
by: Montanari, Giovanni, et al.
Published: (2026)
Adaptive multi-fidelity optimization with fast learning rates
by: Fiegel, Come, et al.
Published: (2026)
by: Fiegel, Come, et al.
Published: (2026)
Covariance-adapting algorithm for semi-bandits with application to sparse rewards
by: Perrault, Pierre, et al.
Published: (2026)
by: Perrault, Pierre, et al.
Published: (2026)
Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions
by: Hait, Soumita, et al.
Published: (2026)
by: Hait, Soumita, et al.
Published: (2026)
Last-Iterate Convergence of Payoff-Based Independent Learning in Zero-Sum Stochastic Games
by: Chen, Zaiwei, et al.
Published: (2024)
by: Chen, Zaiwei, et al.
Published: (2024)
Bandits on graphs and structures
by: Valko, Michal
Published: (2026)
by: Valko, Michal
Published: (2026)
Last-Iterate Convergence: Zero-Sum Games and Constrained Min-Max Optimization
by: Daskalakis, Constantinos, et al.
Published: (2018)
by: Daskalakis, Constantinos, et al.
Published: (2018)
Learning Zero-Sum Linear Quadratic Games with Improved Sample Complexity and Last-Iterate Convergence
by: Wu, Jiduan, et al.
Published: (2023)
by: Wu, Jiduan, et al.
Published: (2023)
Two-Player Zero-Sum Games with Bandit Feedback
by: Yılmaz, Elif, et al.
Published: (2025)
by: Yılmaz, Elif, et al.
Published: (2025)
Learning in Prophet Inequalities with Noisy Observations
by: Kim, Jung-hun, et al.
Published: (2026)
by: Kim, Jung-hun, et al.
Published: (2026)
Strategic Multi-Armed Bandit Problems Under Debt-Free Reporting
by: Yahmed, Ahmed Ben, et al.
Published: (2025)
by: Yahmed, Ahmed Ben, et al.
Published: (2025)
Instance-dependent Stochastic Lipschitz bandit
by: Potfer, Marius, et al.
Published: (2026)
by: Potfer, Marius, et al.
Published: (2026)
On Tradeoffs in Learning-Augmented Algorithms
by: Benomar, Ziyad, et al.
Published: (2025)
by: Benomar, Ziyad, et al.
Published: (2025)
Adaptive Bandit Algorithms for Contextual Matching Markets
by: Lin, Shiyun, et al.
Published: (2026)
by: Lin, Shiyun, et al.
Published: (2026)
Efficient Uncoupled Learning Dynamics with $\tilde{O}\!\left(T^{-1/4}\right)$ Last-Iterate Convergence in Bilinear Saddle-Point Problems over Convex Sets under Bandit Feedback
by: Maiti, Arnab, et al.
Published: (2026)
by: Maiti, Arnab, et al.
Published: (2026)
Online Packet Scheduling with Deadlines and Learning
by: Genalti, Gianmarco, et al.
Published: (2026)
by: Genalti, Gianmarco, et al.
Published: (2026)
The Value of Reward Lookahead in Reinforcement Learning
by: Merlis, Nadav, et al.
Published: (2024)
by: Merlis, Nadav, et al.
Published: (2024)
Unregularized Linear Convergence in Zero-Sum Game from Preference Feedback
by: Chen, Shulun, et al.
Published: (2025)
by: Chen, Shulun, et al.
Published: (2025)
A single algorithm for both restless and rested rotting bandits
by: Seznec, Julien, et al.
Published: (2026)
by: Seznec, Julien, et al.
Published: (2026)
Proximal Point Nash Learning from Human Feedback
by: Tiapkin, Daniil, et al.
Published: (2025)
by: Tiapkin, Daniil, et al.
Published: (2025)
Multi-Armed Bandits with Minimum Aggregated Revenue Constraints
by: Yahmed, Ahmed Ben, et al.
Published: (2025)
by: Yahmed, Ahmed Ben, et al.
Published: (2025)
Comparing Uniform Price and Discriminatory Multi-Unit Auctions through Regret Minimization
by: Potfer, Marius, et al.
Published: (2025)
by: Potfer, Marius, et al.
Published: (2025)
A survey on multi-player bandits
by: Boursier, Etienne, et al.
Published: (2022)
by: Boursier, Etienne, et al.
Published: (2022)
Non-clairvoyant Scheduling with Partial Predictions
by: Benomar, Ziyad, et al.
Published: (2024)
by: Benomar, Ziyad, et al.
Published: (2024)
Bandits attack function optimization
by: Preux, Philippe, et al.
Published: (2026)
by: Preux, Philippe, et al.
Published: (2026)
Mode Estimation with Partial Feedback
by: Arnal, Charles, et al.
Published: (2024)
by: Arnal, Charles, et al.
Published: (2024)
Last-Iterate Convergence of No-Regret Learning for Equilibria in Bargaining Games
by: Kamp, Serafina, et al.
Published: (2025)
by: Kamp, Serafina, et al.
Published: (2025)
Uncoupled and Convergent Learning in Monotone Games under Bandit Feedback
by: Dong, Jing, et al.
Published: (2024)
by: Dong, Jing, et al.
Published: (2024)
Uniform Last-Iterate Guarantee for Bandits and Reinforcement Learning
by: Liu, Junyan, et al.
Published: (2024)
by: Liu, Junyan, et al.
Published: (2024)
On Separation Between Best-Iterate, Random-Iterate, and Last-Iterate Convergence of Learning in Games
by: Cai, Yang, et al.
Published: (2025)
by: Cai, Yang, et al.
Published: (2025)
Fast Last-Iterate Convergence of Learning in Games Requires Forgetful Algorithms
by: Cai, Yang, et al.
Published: (2024)
by: Cai, Yang, et al.
Published: (2024)
Last-Iterate Convergence Properties of Regret-Matching Algorithms in Games
by: Cai, Yang, et al.
Published: (2023)
by: Cai, Yang, et al.
Published: (2023)
Extragradient Preference Optimization (EGPO): Beyond Last-Iterate Convergence for Nash Learning from Human Feedback
by: Zhou, Runlong, et al.
Published: (2025)
by: Zhou, Runlong, et al.
Published: (2025)
Last-Iterate Convergence of Randomized Kaczmarz and SGD with Greedy Step Size
by: Dereziński, Michał, et al.
Published: (2026)
by: Dereziński, Michał, et al.
Published: (2026)
Stable Matching with Ties: Approximation Ratios and Learning
by: Lin, Shiyun, et al.
Published: (2024)
by: Lin, Shiyun, et al.
Published: (2024)
Planning in entropy-regularized Markov decision processes and games
by: Grill, Jean-Bastien, et al.
Published: (2026)
by: Grill, Jean-Bastien, et al.
Published: (2026)
On the Hardness of Reinforcement Learning with Transition Look-Ahead
by: Pla, Corentin, et al.
Published: (2025)
by: Pla, Corentin, et al.
Published: (2025)
Adaptive graph-based algorithms for conditional anomaly detection and semi-supervised learning
by: Valko, Michal
Published: (2026)
by: Valko, Michal
Published: (2026)
Learning Equilibria in Matching Games with Bandit Feedback
by: Athanasopoulos, Andreas, et al.
Published: (2025)
by: Athanasopoulos, Andreas, et al.
Published: (2025)
Similar Items
-
Optimal last-iterate convergence in matrix games with bandit feedback using the log-barrier
by: Fiegel, Come, et al.
Published: (2026) -
Learning to Allocate Resources with Censored Feedback
by: Montanari, Giovanni, et al.
Published: (2026) -
Adaptive multi-fidelity optimization with fast learning rates
by: Fiegel, Come, et al.
Published: (2026) -
Covariance-adapting algorithm for semi-bandits with application to sparse rewards
by: Perrault, Pierre, et al.
Published: (2026) -
Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions
by: Hait, Soumita, et al.
Published: (2026)