GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Costarelli, Anthony, Allen, Mat, Hauksson, Roman, Sodunke, Grace, Hariharan, Suhas, Cheng, Carlson, Li, Wenjie, Clymer, Joshua, Yadav, Arjun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language
di: Costarelli, Anthony, et al.
Pubblicazione: (2024)
di: Costarelli, Anthony, et al.
Pubblicazione: (2024)
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
di: Xu, Zelai, et al.
Pubblicazione: (2025)
di: Xu, Zelai, et al.
Pubblicazione: (2025)
TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
di: Wang, Haochuan, et al.
Pubblicazione: (2024)
di: Wang, Haochuan, et al.
Pubblicazione: (2024)
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
di: Wu, Yunze, et al.
Pubblicazione: (2025)
di: Wu, Yunze, et al.
Pubblicazione: (2025)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
di: Veuthey, Jaime Raldua, et al.
Pubblicazione: (2025)
di: Veuthey, Jaime Raldua, et al.
Pubblicazione: (2025)
Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
di: Hariharan, Suhas, et al.
Pubblicazione: (2024)
di: Hariharan, Suhas, et al.
Pubblicazione: (2024)
DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments
di: Tang, Wenjie, et al.
Pubblicazione: (2025)
di: Tang, Wenjie, et al.
Pubblicazione: (2025)
TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games
di: Mishra, Prakamya, et al.
Pubblicazione: (2025)
di: Mishra, Prakamya, et al.
Pubblicazione: (2025)
Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals
di: Clymer, Joshua, et al.
Pubblicazione: (2024)
di: Clymer, Joshua, et al.
Pubblicazione: (2024)
LocateBench: Evaluating the Locating Ability of Vision Language Models
di: Chiang, Ting-Rui, et al.
Pubblicazione: (2024)
di: Chiang, Ting-Rui, et al.
Pubblicazione: (2024)
Scheming Ability in LLM-to-LLM Strategic Interactions
di: Pham, Thao
Pubblicazione: (2025)
di: Pham, Thao
Pubblicazione: (2025)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
di: Yang, Siwei, et al.
Pubblicazione: (2024)
di: Yang, Siwei, et al.
Pubblicazione: (2024)
Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games
di: He, Yidong, et al.
Pubblicazione: (2026)
di: He, Yidong, et al.
Pubblicazione: (2026)
Study of a longitudinally expanding plasma with the 2PI effective action
di: Gelis, François, et al.
Pubblicazione: (2023)
di: Gelis, François, et al.
Pubblicazione: (2023)
Isotropization of a longitudinally expanding system of scalar fields in the 2PI formalism
di: Gelis, François, et al.
Pubblicazione: (2024)
di: Gelis, François, et al.
Pubblicazione: (2024)
Evaluating Strategic Reasoning in Forecasting Agents
di: Liptay, Tom, et al.
Pubblicazione: (2026)
di: Liptay, Tom, et al.
Pubblicazione: (2026)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
di: Guo, Zhengkang, et al.
Pubblicazione: (2026)
di: Guo, Zhengkang, et al.
Pubblicazione: (2026)
Saturation theorems for neural network operators by solving elliptic and hyperbolic PDEs with analytical and semi-analytical inverse problems
di: Costarelli, Danilo
Pubblicazione: (2025)
di: Costarelli, Danilo
Pubblicazione: (2025)
MAR:Multi-Agent Reflexion Improves Reasoning Abilities in LLMs
di: Ozer, Onat, et al.
Pubblicazione: (2025)
di: Ozer, Onat, et al.
Pubblicazione: (2025)
SpecBench: Evaluating Specification-Level Reasoning for Software Engineering LLM Agents
di: Hamblin, Grant, et al.
Pubblicazione: (2026)
di: Hamblin, Grant, et al.
Pubblicazione: (2026)
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models
di: Yin, Qiyue, et al.
Pubblicazione: (2025)
di: Yin, Qiyue, et al.
Pubblicazione: (2025)
BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
di: Huang, Jiahao, et al.
Pubblicazione: (2026)
di: Huang, Jiahao, et al.
Pubblicazione: (2026)
LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory
di: Jia, Jingru, et al.
Pubblicazione: (2025)
di: Jia, Jingru, et al.
Pubblicazione: (2025)
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
di: Parmar, Mihir, et al.
Pubblicazione: (2024)
di: Parmar, Mihir, et al.
Pubblicazione: (2024)
TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models
di: Chu, Zheng, et al.
Pubblicazione: (2023)
di: Chu, Zheng, et al.
Pubblicazione: (2023)
Safety Cases: How to Justify the Safety of Advanced AI Systems
di: Clymer, Joshua, et al.
Pubblicazione: (2024)
di: Clymer, Joshua, et al.
Pubblicazione: (2024)
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions
di: Cohn, Anthony G, et al.
Pubblicazione: (2024)
di: Cohn, Anthony G, et al.
Pubblicazione: (2024)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
di: Zhang, Zhexin, et al.
Pubblicazione: (2024)
di: Zhang, Zhexin, et al.
Pubblicazione: (2024)
Natural Strategic Ability in Stochastic Multi-Agent Systems
di: Berthon, Raphaël, et al.
Pubblicazione: (2024)
di: Berthon, Raphaël, et al.
Pubblicazione: (2024)
SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
di: Wang, Xiaoxuan, et al.
Pubblicazione: (2023)
di: Wang, Xiaoxuan, et al.
Pubblicazione: (2023)
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
di: Chi, Wayne, et al.
Pubblicazione: (2025)
di: Chi, Wayne, et al.
Pubblicazione: (2025)
JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees
di: Wang, Yuhui, et al.
Pubblicazione: (2026)
di: Wang, Yuhui, et al.
Pubblicazione: (2026)
TPS-Bench: Evaluating AI Agents' Tool Planning \& Scheduling Abilities in Compounding Tasks
di: Xu, Hanwen, et al.
Pubblicazione: (2025)
di: Xu, Hanwen, et al.
Pubblicazione: (2025)
Reasoning about Strategic Abilities in Stochastic Multi-agent Systems
di: Zhang, Yedi, et al.
Pubblicazione: (2024)
di: Zhang, Yedi, et al.
Pubblicazione: (2024)
María Agustina Juri y Fernando Martín De Blassi, Alberto Magno: Las Virtudes Cardinales (In Sent. Iii, D. 33). Edición Bilingüe, Introducción, Traducción y Notas. Ápeiron, Madrid, 2018, 105 PP. ISBN 978-84-17574-68-0
di: Hugo Costarelli Brandi
Pubblicazione: (2019)
di: Hugo Costarelli Brandi
Pubblicazione: (2019)
Naturaleza y ficción en la imitación artística: consideraciones desde Aristóteles
di: Hugo Costarelli Brandi
Pubblicazione: (2023)
di: Hugo Costarelli Brandi
Pubblicazione: (2023)
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited
di: Cohn, Anthony G, et al.
Pubblicazione: (2025)
di: Cohn, Anthony G, et al.
Pubblicazione: (2025)
Enhancing Language Agent Strategic Reasoning through Self-Play in Adversarial Games
di: Zhang, Yikai, et al.
Pubblicazione: (2025)
di: Zhang, Yikai, et al.
Pubblicazione: (2025)
Dataset Featurization: Uncovering Natural Language Features through Unsupervised Data Reconstruction
di: Bravansky, Michal, et al.
Pubblicazione: (2025)
di: Bravansky, Michal, et al.
Pubblicazione: (2025)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
di: Maniparambil, Mayug, et al.
Pubblicazione: (2026)
di: Maniparambil, Mayug, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Meta-Models: An Architecture for Decoding LLM Behaviors Through Interpreted Embeddings and Natural Language
di: Costarelli, Anthony, et al.
Pubblicazione: (2024) -
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
di: Xu, Zelai, et al.
Pubblicazione: (2025) -
TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs
di: Wang, Haochuan, et al.
Pubblicazione: (2024) -
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
di: Wu, Yunze, et al.
Pubblicazione: (2025) -
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
di: Veuthey, Jaime Raldua, et al.
Pubblicazione: (2025)