GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Costarelli, Anthony, Allen, Mat, Hauksson, Roman, Sodunke, Grace, Hariharan, Suhas, Cheng, Carlson, Li, Wenjie, Clymer, Joshua, Yadav, Arjun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language
by: Costarelli, Anthony, et al.
Published: (2024)
by: Costarelli, Anthony, et al.
Published: (2024)
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
by: Xu, Zelai, et al.
Published: (2025)
by: Xu, Zelai, et al.
Published: (2025)
TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
by: Wang, Haochuan, et al.
Published: (2024)
by: Wang, Haochuan, et al.
Published: (2024)
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
by: Wu, Yunze, et al.
Published: (2025)
by: Wu, Yunze, et al.
Published: (2025)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
by: Hariharan, Suhas, et al.
Published: (2024)
by: Hariharan, Suhas, et al.
Published: (2024)
DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments
by: Tang, Wenjie, et al.
Published: (2025)
by: Tang, Wenjie, et al.
Published: (2025)
TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games
by: Mishra, Prakamya, et al.
Published: (2025)
by: Mishra, Prakamya, et al.
Published: (2025)
Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals
by: Clymer, Joshua, et al.
Published: (2024)
by: Clymer, Joshua, et al.
Published: (2024)
LocateBench: Evaluating the Locating Ability of Vision Language Models
by: Chiang, Ting-Rui, et al.
Published: (2024)
by: Chiang, Ting-Rui, et al.
Published: (2024)
Scheming Ability in LLM-to-LLM Strategic Interactions
by: Pham, Thao
Published: (2025)
by: Pham, Thao
Published: (2025)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
by: Yang, Siwei, et al.
Published: (2024)
by: Yang, Siwei, et al.
Published: (2024)
Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games
by: He, Yidong, et al.
Published: (2026)
by: He, Yidong, et al.
Published: (2026)
Study of a longitudinally expanding plasma with the 2PI effective action
by: Gelis, François, et al.
Published: (2023)
by: Gelis, François, et al.
Published: (2023)
Isotropization of a longitudinally expanding system of scalar fields in the 2PI formalism
by: Gelis, François, et al.
Published: (2024)
by: Gelis, François, et al.
Published: (2024)
Evaluating Strategic Reasoning in Forecasting Agents
by: Liptay, Tom, et al.
Published: (2026)
by: Liptay, Tom, et al.
Published: (2026)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
by: Guo, Zhengkang, et al.
Published: (2026)
by: Guo, Zhengkang, et al.
Published: (2026)
Saturation theorems for neural network operators by solving elliptic and hyperbolic PDEs with analytical and semi-analytical inverse problems
by: Costarelli, Danilo
Published: (2025)
by: Costarelli, Danilo
Published: (2025)
MAR:Multi-Agent Reflexion Improves Reasoning Abilities in LLMs
by: Ozer, Onat, et al.
Published: (2025)
by: Ozer, Onat, et al.
Published: (2025)
SpecBench: Evaluating Specification-Level Reasoning for Software Engineering LLM Agents
by: Hamblin, Grant, et al.
Published: (2026)
by: Hamblin, Grant, et al.
Published: (2026)
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models
by: Yin, Qiyue, et al.
Published: (2025)
by: Yin, Qiyue, et al.
Published: (2025)
BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
by: Huang, Jiahao, et al.
Published: (2026)
by: Huang, Jiahao, et al.
Published: (2026)
LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory
by: Jia, Jingru, et al.
Published: (2025)
by: Jia, Jingru, et al.
Published: (2025)
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
by: Parmar, Mihir, et al.
Published: (2024)
by: Parmar, Mihir, et al.
Published: (2024)
TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models
by: Chu, Zheng, et al.
Published: (2023)
by: Chu, Zheng, et al.
Published: (2023)
Safety Cases: How to Justify the Safety of Advanced AI Systems
by: Clymer, Joshua, et al.
Published: (2024)
by: Clymer, Joshua, et al.
Published: (2024)
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions
by: Cohn, Anthony G, et al.
Published: (2024)
by: Cohn, Anthony G, et al.
Published: (2024)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
by: Zhang, Zhexin, et al.
Published: (2024)
by: Zhang, Zhexin, et al.
Published: (2024)
Natural Strategic Ability in Stochastic Multi-Agent Systems
by: Berthon, Raphaël, et al.
Published: (2024)
by: Berthon, Raphaël, et al.
Published: (2024)
SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
by: Wang, Xiaoxuan, et al.
Published: (2023)
by: Wang, Xiaoxuan, et al.
Published: (2023)
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
by: Chi, Wayne, et al.
Published: (2025)
by: Chi, Wayne, et al.
Published: (2025)
JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees
by: Wang, Yuhui, et al.
Published: (2026)
by: Wang, Yuhui, et al.
Published: (2026)
TPS-Bench: Evaluating AI Agents' Tool Planning \& Scheduling Abilities in Compounding Tasks
by: Xu, Hanwen, et al.
Published: (2025)
by: Xu, Hanwen, et al.
Published: (2025)
Reasoning about Strategic Abilities in Stochastic Multi-agent Systems
by: Zhang, Yedi, et al.
Published: (2024)
by: Zhang, Yedi, et al.
Published: (2024)
María Agustina Juri y Fernando Martín De Blassi, Alberto Magno: Las Virtudes Cardinales (In Sent. Iii, D. 33). Edición Bilingüe, Introducción, Traducción y Notas. Ápeiron, Madrid, 2018, 105 PP. ISBN 978-84-17574-68-0
by: Hugo Costarelli Brandi
Published: (2019)
by: Hugo Costarelli Brandi
Published: (2019)
Naturaleza y ficción en la imitación artística: consideraciones desde Aristóteles
by: Hugo Costarelli Brandi
Published: (2023)
by: Hugo Costarelli Brandi
Published: (2023)
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited
by: Cohn, Anthony G, et al.
Published: (2025)
by: Cohn, Anthony G, et al.
Published: (2025)
Enhancing Language Agent Strategic Reasoning through Self-Play in Adversarial Games
by: Zhang, Yikai, et al.
Published: (2025)
by: Zhang, Yikai, et al.
Published: (2025)
Dataset Featurization: Uncovering Natural Language Features through Unsupervised Data Reconstruction
by: Bravansky, Michal, et al.
Published: (2025)
by: Bravansky, Michal, et al.
Published: (2025)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
by: Maniparambil, Mayug, et al.
Published: (2026)
by: Maniparambil, Mayug, et al.
Published: (2026)
Similar Items
-
Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language
by: Costarelli, Anthony, et al.
Published: (2024) -
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
by: Xu, Zelai, et al.
Published: (2025) -
TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
by: Wang, Haochuan, et al.
Published: (2024) -
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
by: Wu, Yunze, et al.
Published: (2025) -
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
by: Veuthey, Jaime Raldua, et al.
Published: (2025)