Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.10673 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908323064840192 |
|---|---|
| author | Alyahya, Hisham A. Khan, Haidar Alnumay, Yazeed Bari, M Saiful Yener, Bülent |
| author_facet | Alyahya, Hisham A. Khan, Haidar Alnumay, Yazeed Bari, M Saiful Yener, Bülent |
| contents | We introduce ZeroSumEval, a dynamic, competition-based, and evolving evaluation framework for Large Language Models (LLMs) that leverages competitive games. ZeroSumEval encompasses a diverse suite of games, including security challenges (Capture the Flag), classic board games (chess), and knowledge tests (MathQuiz). These games are designed to evaluate a range of capabilities such as strategic reasoning, planning, knowledge application, safety, and adaptability. Building upon recent studies that highlight the effectiveness of game-based evaluations for LLMs, ZeroSumEval enhances these approaches by providing a standardized and extensible framework for easily implementing games and leverages DSPy to provide a better abstraction for LLM player strategies. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_10673 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition Alyahya, Hisham A. Khan, Haidar Alnumay, Yazeed Bari, M Saiful Yener, Bülent Computation and Language Artificial Intelligence We introduce ZeroSumEval, a dynamic, competition-based, and evolving evaluation framework for Large Language Models (LLMs) that leverages competitive games. ZeroSumEval encompasses a diverse suite of games, including security challenges (Capture the Flag), classic board games (chess), and knowledge tests (MathQuiz). These games are designed to evaluate a range of capabilities such as strategic reasoning, planning, knowledge application, safety, and adaptability. Building upon recent studies that highlight the effectiveness of game-based evaluations for LLMs, ZeroSumEval enhances these approaches by providing a standardized and extensible framework for easily implementing games and leverages DSPy to provide a better abstraction for LLM player strategies. |
| title | ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2503.10673 |