Saved in:
Bibliographic Details
Main Authors: Khan, Haidar, Alyahya, Hisham A., Alnumay, Yazeed, Bari, M Saiful, Yener, Bülent
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.12562
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917987693363200
author Khan, Haidar
Alyahya, Hisham A.
Alnumay, Yazeed
Bari, M Saiful
Yener, Bülent
author_facet Khan, Haidar
Alyahya, Hisham A.
Alnumay, Yazeed
Bari, M Saiful
Yener, Bülent
contents Evaluating the capabilities of Large Language Models (LLMs) has traditionally relied on static benchmark datasets, human assessments, or model-based evaluations - methods that often suffer from overfitting, high costs, and biases. ZeroSumEval is a novel competition-based evaluation protocol that leverages zero-sum games to assess LLMs with dynamic benchmarks that resist saturation. ZeroSumEval encompasses a diverse suite of games, including security challenges (PyJail), classic games (Chess, Liar's Dice, Poker), knowledge tests (MathQuiz), and persuasion challenges (Gandalf, Debate). These games are designed to evaluate a range of AI capabilities such as strategic reasoning, planning, knowledge application, and creativity. Building upon recent studies that highlight the effectiveness of game-based evaluations for LLMs, ZeroSumEval enhances these approaches by providing a standardized and extensible framework. To demonstrate this, we conduct extensive experiments with >7000 simulations across 7 games and 13 models. Our results show that while frontier models from the GPT and Claude families can play common games and answer questions, they struggle to play games that require creating novel and challenging questions. We also observe that models cannot reliably jailbreak each other and fail generally at tasks requiring creativity. We release our code at https://github.com/facebookresearch/ZeroSumEval.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12562
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
Khan, Haidar
Alyahya, Hisham A.
Alnumay, Yazeed
Bari, M Saiful
Yener, Bülent
Artificial Intelligence
Computation and Language
Evaluating the capabilities of Large Language Models (LLMs) has traditionally relied on static benchmark datasets, human assessments, or model-based evaluations - methods that often suffer from overfitting, high costs, and biases. ZeroSumEval is a novel competition-based evaluation protocol that leverages zero-sum games to assess LLMs with dynamic benchmarks that resist saturation. ZeroSumEval encompasses a diverse suite of games, including security challenges (PyJail), classic games (Chess, Liar's Dice, Poker), knowledge tests (MathQuiz), and persuasion challenges (Gandalf, Debate). These games are designed to evaluate a range of AI capabilities such as strategic reasoning, planning, knowledge application, and creativity. Building upon recent studies that highlight the effectiveness of game-based evaluations for LLMs, ZeroSumEval enhances these approaches by providing a standardized and extensible framework. To demonstrate this, we conduct extensive experiments with >7000 simulations across 7 games and 13 models. Our results show that while frontier models from the GPT and Claude families can play common games and answer questions, they struggle to play games that require creating novel and challenging questions. We also observe that models cannot reliably jailbreak each other and fail generally at tasks requiring creativity. We release our code at https://github.com/facebookresearch/ZeroSumEval.
title ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2504.12562