Re-evaluating Open-ended Evaluation of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Siqi, Gemp, Ian, Marris, Luke, Piliouras, Georgios, Heess, Nicolas, Lanctot, Marc |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deviation Ratings: A General, Clone-Invariant Rating Method
by: Marris, Luke, et al.
Published: (2025)
by: Marris, Luke, et al.
Published: (2025)
NfgTransformer: Equivariant Representation Learning for Normal-form Games
by: Liu, Siqi, et al.
Published: (2024)
by: Liu, Siqi, et al.
Published: (2024)
Steering Language Models with Game-Theoretic Solvers
by: Gemp, Ian, et al.
Published: (2024)
by: Gemp, Ian, et al.
Published: (2024)
Approximating Nash Equilibria in Normal-Form Games via Stochastic Optimization
by: Gemp, Ian, et al.
Published: (2023)
by: Gemp, Ian, et al.
Published: (2023)
Visualizing 2x2 Normal-Form Games: twoxtwogame LaTeX Package
by: Marris, Luke, et al.
Published: (2024)
by: Marris, Luke, et al.
Published: (2024)
Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning
by: Gemp, Ian, et al.
Published: (2024)
by: Gemp, Ian, et al.
Published: (2024)
Deep Incentive Design with Differentiable Equilibrium Blocks
by: Thoma, Vinzenz, et al.
Published: (2026)
by: Thoma, Vinzenz, et al.
Published: (2026)
Tight Inapproximability for Welfare-Maximizing Autobidding Equilibria
by: Anagnostides, Ioannis, et al.
Published: (2026)
by: Anagnostides, Ioannis, et al.
Published: (2026)
Chaos in Autobidding Auctions
by: Anagnostides, Ioannis, et al.
Published: (2026)
by: Anagnostides, Ioannis, et al.
Published: (2026)
Approximating the Core via Iterative Coalition Sampling
by: Gemp, Ian, et al.
Published: (2024)
by: Gemp, Ian, et al.
Published: (2024)
Active Evaluation of General Agents: Problem Definition and Comparison of Baseline Algorithms
by: Lanctot, Marc, et al.
Published: (2026)
by: Lanctot, Marc, et al.
Published: (2026)
Nash without Numbers: A Social Choice Approach to Mixed Equilibria in Context-Ordinal Games
by: Gemp, Ian, et al.
Published: (2026)
by: Gemp, Ian, et al.
Published: (2026)
Charting the Shapes of Stories with Game Theory
by: Daskalakis, Constantinos, et al.
Published: (2024)
by: Daskalakis, Constantinos, et al.
Published: (2024)
Nash Equilibria via Stochastic Eigendecomposition
by: Gemp, Ian
Published: (2024)
by: Gemp, Ian
Published: (2024)
Solving Zero-Sum Convex Markov Games
by: Kalogiannis, Fivos, et al.
Published: (2025)
by: Kalogiannis, Fivos, et al.
Published: (2025)
Multi-Agent Training beyond Zero-Sum with Correlated Equilibrium Meta-Solvers
by: Marris, Luke, et al.
Published: (2021)
by: Marris, Luke, et al.
Published: (2021)
Combining Tree-Search, Generative Models, and Nash Bargaining Concepts in Game-Theoretic Reinforcement Learning
by: Li, Zun, et al.
Published: (2023)
by: Li, Zun, et al.
Published: (2023)
Prediction Accuracy of Learning in Games : Follow-the-Regularized-Leader meets Heisenberg
by: Feng, Yi, et al.
Published: (2024)
by: Feng, Yi, et al.
Published: (2024)
Learning in Games with Progressive Hiding
by: Heymann, Benjamin, et al.
Published: (2024)
by: Heymann, Benjamin, et al.
Published: (2024)
On the Complexity of Learning Nash Equilibria
by: Biggar, Oliver, et al.
Published: (2026)
by: Biggar, Oliver, et al.
Published: (2026)
Evaluating Agents using Social Choice Theory
by: Lanctot, Marc, et al.
Published: (2023)
by: Lanctot, Marc, et al.
Published: (2023)
Discovering Multiagent Learning Algorithms with Large Language Models
by: Li, Zun, et al.
Published: (2026)
by: Li, Zun, et al.
Published: (2026)
Eliciting Informative Text Evaluations with Large Language Models
by: Lu, Yuxuan, et al.
Published: (2024)
by: Lu, Yuxuan, et al.
Published: (2024)
Fairshare Data Pricing via Data Valuation for Large Language Models
by: Zhang, Luyang, et al.
Published: (2025)
by: Zhang, Luyang, et al.
Published: (2025)
Complex Dynamics in Autobidding Systems
by: Leme, Renato Paes, et al.
Published: (2024)
by: Leme, Renato Paes, et al.
Published: (2024)
Global Convergence of Multi-Agent Policy Gradient in Markov Potential Games
by: Leonardos, Stefanos, et al.
Published: (2021)
by: Leonardos, Stefanos, et al.
Published: (2021)
Code-Space Response Oracles: Generating Interpretable Multi-Agent Policies with Large Language Models
by: Hennes, Daniel, et al.
Published: (2026)
by: Hennes, Daniel, et al.
Published: (2026)
Fast and Furious Symmetric Learning in Zero-Sum Games: Gradient Descent as Fictitious Play
by: Lazarsfeld, John, et al.
Published: (2025)
by: Lazarsfeld, John, et al.
Published: (2025)
No-Regret Learning and Equilibrium Computation in Quantum Games
by: Lin, Wayne, et al.
Published: (2023)
by: Lin, Wayne, et al.
Published: (2023)
Learning in Quantum Common-Interest Games and the Separability Problem
by: Lin, Wayne, et al.
Published: (2023)
by: Lin, Wayne, et al.
Published: (2023)
Optimism Without Regularization: Constant Regret in Zero-Sum Games
by: Lazarsfeld, John, et al.
Published: (2025)
by: Lazarsfeld, John, et al.
Published: (2025)
Faster Rates for No-Regret Learning in General Games via Cautious Optimism
by: Soleymani, Ashkan, et al.
Published: (2025)
by: Soleymani, Ashkan, et al.
Published: (2025)
Cautious Optimism: A Meta-Algorithm for Near-Constant Regret in General Games
by: Soleymani, Ashkan, et al.
Published: (2025)
by: Soleymani, Ashkan, et al.
Published: (2025)
Data-Scarce Identification of Game Dynamics via Sum-of-Squares Optimization
by: Sakos, Iosif, et al.
Published: (2023)
by: Sakos, Iosif, et al.
Published: (2023)
When and Why is Optimistic Multiplicative Weights Slow? The Geometry of Energy Dissipation
by: Lazarsfeld, John, et al.
Published: (2026)
by: Lazarsfeld, John, et al.
Published: (2026)
Learning and steering game dynamics towards desirable outcomes
by: Canyakmaz, Ilayda, et al.
Published: (2024)
by: Canyakmaz, Ilayda, et al.
Published: (2024)
Strange bifurcation diagrams
by: Bielawski, Jakub, et al.
Published: (2026)
by: Bielawski, Jakub, et al.
Published: (2026)
Evolving Diverse Red-team Language Models in Multi-round Multi-agent Games
by: Ma, Chengdong, et al.
Published: (2023)
by: Ma, Chengdong, et al.
Published: (2023)
Distributive Fairness in Large Language Models: Evaluating Alignment with Human Values
by: Hosseini, Hadi, et al.
Published: (2025)
by: Hosseini, Hadi, et al.
Published: (2025)
PokerBench: Training Large Language Models to become Professional Poker Players
by: Zhuang, Richard, et al.
Published: (2025)
by: Zhuang, Richard, et al.
Published: (2025)
Similar Items
-
Deviation Ratings: A General, Clone-Invariant Rating Method
by: Marris, Luke, et al.
Published: (2025) -
NfgTransformer: Equivariant Representation Learning for Normal-form Games
by: Liu, Siqi, et al.
Published: (2024) -
Steering Language Models with Game-Theoretic Solvers
by: Gemp, Ian, et al.
Published: (2024) -
Approximating Nash Equilibria in Normal-Form Games via Stochastic Optimization
by: Gemp, Ian, et al.
Published: (2023) -
Visualizing 2x2 Normal-Form Games: twoxtwogame LaTeX Package
by: Marris, Luke, et al.
Published: (2024)