Online Learning and Equilibrium Computation with Ranking Feedback
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Mingyang, Chen, Yongshan, Fan, Zhiyuan, Farina, Gabriele, Ozdaglar, Asuman, Zhang, Kaiqing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Computing Equilibrium beyond Unilateral Deviation
por: Liu, Mingyang, et al.
Publicado: (2026)
por: Liu, Mingyang, et al.
Publicado: (2026)
Differentially Private Equilibrium Finding in Polymatrix Games
por: Liu, Mingyang, et al.
Publicado: (2025)
por: Liu, Mingyang, et al.
Publicado: (2025)
LiteEFG: An Efficient Python Library for Solving Extensive-form Games
por: Liu, Mingyang, et al.
Publicado: (2024)
por: Liu, Mingyang, et al.
Publicado: (2024)
A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence
por: Liu, Mingyang, et al.
Publicado: (2024)
por: Liu, Mingyang, et al.
Publicado: (2024)
The Power of Regularization in Solving Extensive-Form Games
por: Liu, Mingyang, et al.
Publicado: (2022)
por: Liu, Mingyang, et al.
Publicado: (2022)
Multi-Player Zero-Sum Markov Games with Networked Separable Interactions
por: Park, Chanwoo, et al.
Publicado: (2023)
por: Park, Chanwoo, et al.
Publicado: (2023)
Do LLM Agents Have Regret? A Case Study in Online Learning and Games
por: Park, Chanwoo, et al.
Publicado: (2024)
por: Park, Chanwoo, et al.
Publicado: (2024)
Last-Iterate Convergence of Payoff-Based Independent Learning in Zero-Sum Stochastic Games
por: Chen, Zaiwei, et al.
Publicado: (2024)
por: Chen, Zaiwei, et al.
Publicado: (2024)
On the Optimality of Dilated Entropy and Lower Bounds for Online Learning in Extensive-Form Games
por: Fan, Zhiyuan, et al.
Publicado: (2024)
por: Fan, Zhiyuan, et al.
Publicado: (2024)
GAE Falls Short in Imperfect-Information Self-Play Reinforcement Learning
por: Fan, Zhiyuan, et al.
Publicado: (2026)
por: Fan, Zhiyuan, et al.
Publicado: (2026)
Efficient Near-Optimal Algorithm for Online Shortest Paths in Directed Acyclic Graphs with Bandit Feedback Against Adaptive Adversaries
por: Maiti, Arnab, et al.
Publicado: (2025)
por: Maiti, Arnab, et al.
Publicado: (2025)
Finite-Sample Guarantees for Learning Dynamics in Zero-Sum Polymatrix Games
por: Faizal, Fathima Zarin, et al.
Publicado: (2024)
por: Faizal, Fathima Zarin, et al.
Publicado: (2024)
UFT: Unifying Supervised and Reinforcement Fine-Tuning
por: Liu, Mingyang, et al.
Publicado: (2025)
por: Liu, Mingyang, et al.
Publicado: (2025)
An Efficient Black-Box Reduction from Online Learning to Multicalibration, and a New Route to $Φ$-Regret Minimization
por: Farina, Gabriele, et al.
Publicado: (2026)
por: Farina, Gabriele, et al.
Publicado: (2026)
Equilibrium Selection for Multi-agent Reinforcement Learning: A Unified Framework
por: Zhang, Runyu, et al.
Publicado: (2024)
por: Zhang, Runyu, et al.
Publicado: (2024)
The Stability of Online Algorithms in Performative Prediction
por: Farina, Gabriele, et al.
Publicado: (2026)
por: Farina, Gabriele, et al.
Publicado: (2026)
On the Universal Near Optimality of Hedge in Combinatorial Settings
por: Fan, Zhiyuan, et al.
Publicado: (2025)
por: Fan, Zhiyuan, et al.
Publicado: (2025)
Partially Observable Multi-Agent Reinforcement Learning with Information Sharing
por: Liu, Xiangyu, et al.
Publicado: (2023)
por: Liu, Xiangyu, et al.
Publicado: (2023)
Learning and Computation of $Φ$-Equilibria at the Frontier of Tractability
por: Zhang, Brian Hu, et al.
Publicado: (2025)
por: Zhang, Brian Hu, et al.
Publicado: (2025)
Online Learning for Equilibrium Pricing in Markets under Incomplete Information
por: Jalota, Devansh, et al.
Publicado: (2023)
por: Jalota, Devansh, et al.
Publicado: (2023)
Efficient Learning and Computation of Linear Correlated Equilibrium in General Convex Games
por: Daskalakis, Constantinos, et al.
Publicado: (2024)
por: Daskalakis, Constantinos, et al.
Publicado: (2024)
Hidden-Role Games: Equilibrium Concepts and Computation
por: Carminati, Luca, et al.
Publicado: (2023)
por: Carminati, Luca, et al.
Publicado: (2023)
Nash CoT: Multi-Path Inference with Preference Equilibrium
por: Zhang, Ziqi, et al.
Publicado: (2024)
por: Zhang, Ziqi, et al.
Publicado: (2024)
Faster Rates for No-Regret Learning in General Games via Cautious Optimism
por: Soleymani, Ashkan, et al.
Publicado: (2025)
por: Soleymani, Ashkan, et al.
Publicado: (2025)
Matching of Users and Creators in Two-Sided Markets with Departures
por: Huttenlocher, Daniel, et al.
Publicado: (2023)
por: Huttenlocher, Daniel, et al.
Publicado: (2023)
Online Budget Allocation with Censored Semi-Bandit Feedback
por: Bachoc, François, et al.
Publicado: (2025)
por: Bachoc, François, et al.
Publicado: (2025)
Large-Scale Contextual Market Equilibrium Computation through Deep Learning
por: Ma, Yunxuan, et al.
Publicado: (2024)
por: Ma, Yunxuan, et al.
Publicado: (2024)
Cautious Optimism: A Meta-Algorithm for Near-Constant Regret in General Games
por: Soleymani, Ashkan, et al.
Publicado: (2025)
por: Soleymani, Ashkan, et al.
Publicado: (2025)
Optimal Correlated Equilibria in General-Sum Extensive-Form Games: Fixed-Parameter Algorithms, Hardness, and Two-Sided Column-Generation
por: Zhang, Brian, et al.
Publicado: (2022)
por: Zhang, Brian, et al.
Publicado: (2022)
Taming Equilibrium Bias in Risk-Sensitive Multi-Agent Reinforcement Learning
por: Fei, Yingjie, et al.
Publicado: (2024)
por: Fei, Yingjie, et al.
Publicado: (2024)
Learning to Allocate Resources with Censored Feedback
por: Montanari, Giovanni, et al.
Publicado: (2026)
por: Montanari, Giovanni, et al.
Publicado: (2026)
Doubly Optimal No-Regret Online Learning in Strongly Monotone Games with Bandit Feedback
por: Ba, Wenjia, et al.
Publicado: (2021)
por: Ba, Wenjia, et al.
Publicado: (2021)
Leaderboard Incentives: Model Rankings under Strategic Post-Training
por: Chen, Yatong, et al.
Publicado: (2026)
por: Chen, Yatong, et al.
Publicado: (2026)
VickreyFeedback: Cost-efficient Data Construction for Reinforcement Learning from Human Feedback
por: Zhang, Guoxi, et al.
Publicado: (2024)
por: Zhang, Guoxi, et al.
Publicado: (2024)
Re-evaluating Open-ended Evaluation of Large Language Models
por: Liu, Siqi, et al.
Publicado: (2025)
por: Liu, Siqi, et al.
Publicado: (2025)
Polynomial-Time Computation of Exact $Φ$-Equilibria in Polyhedral Games
por: Farina, Gabriele, et al.
Publicado: (2024)
por: Farina, Gabriele, et al.
Publicado: (2024)
Improved Regret Bounds for Online Fair Division with Bandit Learning
por: Schiffer, Benjamin, et al.
Publicado: (2025)
por: Schiffer, Benjamin, et al.
Publicado: (2025)
A Polynomial-Time Algorithm for Variational Inequalities under the Minty Condition
por: Anagnostides, Ioannis, et al.
Publicado: (2025)
por: Anagnostides, Ioannis, et al.
Publicado: (2025)
Online Learning for Uninformed Markov Games: Empirical Nash-Value Regret and Non-Stationarity Adaptation
por: Liu, Junyan, et al.
Publicado: (2026)
por: Liu, Junyan, et al.
Publicado: (2026)
Last-Iterate Convergence Properties of Regret-Matching Algorithms in Games
por: Cai, Yang, et al.
Publicado: (2023)
por: Cai, Yang, et al.
Publicado: (2023)
Ejemplares similares
-
Computing Equilibrium beyond Unilateral Deviation
por: Liu, Mingyang, et al.
Publicado: (2026) -
Differentially Private Equilibrium Finding in Polymatrix Games
por: Liu, Mingyang, et al.
Publicado: (2025) -
LiteEFG: An Efficient Python Library for Solving Extensive-form Games
por: Liu, Mingyang, et al.
Publicado: (2024) -
A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence
por: Liu, Mingyang, et al.
Publicado: (2024) -
The Power of Regularization in Solving Extensive-Form Games
por: Liu, Mingyang, et al.
Publicado: (2022)