Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play
Fuente:
arXiv
Salvato in:
| Autore principale: | Ekne, H. C. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Large Language Models as Pokémon Battle Agents: Strategic Play and Content Generation
di: Jain, Daksh, et al.
Pubblicazione: (2025)
di: Jain, Daksh, et al.
Pubblicazione: (2025)
A Multi-Agent Pokemon Tournament for Evaluating Strategic Reasoning of Large Language Models
di: Yashwanth, Tadisetty Sai, et al.
Pubblicazione: (2025)
di: Yashwanth, Tadisetty Sai, et al.
Pubblicazione: (2025)
Role-Playing Evaluation for Large Language Models
di: Boudouri, Yassine El, et al.
Pubblicazione: (2025)
di: Boudouri, Yassine El, et al.
Pubblicazione: (2025)
Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game
di: Xu, Zelai, et al.
Pubblicazione: (2023)
di: Xu, Zelai, et al.
Pubblicazione: (2023)
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
di: Wang, Kevin, et al.
Pubblicazione: (2026)
di: Wang, Kevin, et al.
Pubblicazione: (2026)
Three Roles, One Model: Role Orchestration at Inference Time to Close the Performance Gap Between Small and Large Agents
di: McClendon, S. Aaron, et al.
Pubblicazione: (2026)
di: McClendon, S. Aaron, et al.
Pubblicazione: (2026)
MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
di: Yuan, Huining, et al.
Pubblicazione: (2025)
di: Yuan, Huining, et al.
Pubblicazione: (2025)
LLM-Gomoku: A Large Language Model-Based System for Strategic Gomoku with Self-Play and Reinforcement Learning
di: Wang, Hui
Pubblicazione: (2025)
di: Wang, Hui
Pubblicazione: (2025)
RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines
di: Yu, Pengfei, et al.
Pubblicazione: (2025)
di: Yu, Pengfei, et al.
Pubblicazione: (2025)
Plug-and-Play Policy Planner for Large Language Model Powered Dialogue Agents
di: Deng, Yang, et al.
Pubblicazione: (2023)
di: Deng, Yang, et al.
Pubblicazione: (2023)
Collaborative Intelligence: Topic Modelling of Large Language Model use in Live Cybersecurity Operations
di: Lochner, Martin, et al.
Pubblicazione: (2025)
di: Lochner, Martin, et al.
Pubblicazione: (2025)
Evaluating Strategic Reasoning in Forecasting Agents
di: Liptay, Tom, et al.
Pubblicazione: (2026)
di: Liptay, Tom, et al.
Pubblicazione: (2026)
Can Large Language Models Bridge the Gap in Environmental Knowledge?
di: Smail, Linda, et al.
Pubblicazione: (2025)
di: Smail, Linda, et al.
Pubblicazione: (2025)
A Proposal for Evaluating the Operational Risk for ChatBots based on Large Language Models
di: Pinacho-Davidson, Pedro, et al.
Pubblicazione: (2025)
di: Pinacho-Davidson, Pedro, et al.
Pubblicazione: (2025)
Modular Task Decomposition and Dynamic Collaboration in Multi-Agent Systems Driven by Large Language Models
di: Pan, Shuaidong, et al.
Pubblicazione: (2025)
di: Pan, Shuaidong, et al.
Pubblicazione: (2025)
JADE: Bridging the Strategic-Operational Gap in Dynamic Agentic RAG
di: Chen, Yiqun, et al.
Pubblicazione: (2026)
di: Chen, Yiqun, et al.
Pubblicazione: (2026)
PersonaArena: Dynamic Simulation for Evaluating and Enhancing Persona-Level Role-Playing in Large Language Models
di: Shi, Wenlong, et al.
Pubblicazione: (2026)
di: Shi, Wenlong, et al.
Pubblicazione: (2026)
Strategic Data Ordering: Enhancing Large Language Model Performance through Curriculum Learning
di: Kim, Jisu, et al.
Pubblicazione: (2024)
di: Kim, Jisu, et al.
Pubblicazione: (2024)
Bridging the Evaluation Gap: Leveraging Large Language Models for Topic Model Evaluation
di: Tan, Zhiyin, et al.
Pubblicazione: (2025)
di: Tan, Zhiyin, et al.
Pubblicazione: (2025)
Can Large Language Models Play Games? A Case Study of A Self-Play Approach
di: Guo, Hongyi, et al.
Pubblicazione: (2024)
di: Guo, Hongyi, et al.
Pubblicazione: (2024)
A Survey on Game Playing Agents and Large Models: Methods, Applications, and Challenges
di: Xu, Xinrun, et al.
Pubblicazione: (2024)
di: Xu, Xinrun, et al.
Pubblicazione: (2024)
INSIGHT: Bridging the Student-Teacher Gap in Times of Large Language Models
di: Thys, Jarne, et al.
Pubblicazione: (2025)
di: Thys, Jarne, et al.
Pubblicazione: (2025)
Performance Evaluation of Large Language Models in Statistical Programming
di: Song, Xinyi, et al.
Pubblicazione: (2025)
di: Song, Xinyi, et al.
Pubblicazione: (2025)
Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia
di: Costa, Davi Bastos, et al.
Pubblicazione: (2025)
di: Costa, Davi Bastos, et al.
Pubblicazione: (2025)
Hypergame Rationalisability: Solving Agent Misalignment In Strategic Play
di: Trencsenyi, Vince
Pubblicazione: (2025)
di: Trencsenyi, Vince
Pubblicazione: (2025)
Fairness Testing of Large Language Models in Role-Playing
di: Li, Xinyue, et al.
Pubblicazione: (2024)
di: Li, Xinyue, et al.
Pubblicazione: (2024)
Show, Don't Tell: Evaluating Large Language Models Beyond Textual Understanding with ChildPlay
di: de Carvalho, Gonçalo Hora, et al.
Pubblicazione: (2024)
di: de Carvalho, Gonçalo Hora, et al.
Pubblicazione: (2024)
GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games
di: Li, Yuchen, et al.
Pubblicazione: (2025)
di: Li, Yuchen, et al.
Pubblicazione: (2025)
Unplug and Play Language Models: Decomposing Experts in Language Models at Inference Time
di: Yang, Nakyeong, et al.
Pubblicazione: (2024)
di: Yang, Nakyeong, et al.
Pubblicazione: (2024)
Development and Validation of the Provider Documentation Summarization Quality Instrument for Large Language Models
di: Croxford, Emma, et al.
Pubblicazione: (2025)
di: Croxford, Emma, et al.
Pubblicazione: (2025)
CHATATC: Large Language Model-Driven Conversational Agents for Supporting Strategic Air Traffic Flow Management
di: Abdulhak, Sinan, et al.
Pubblicazione: (2024)
di: Abdulhak, Sinan, et al.
Pubblicazione: (2024)
ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models
di: Liu, Jincheng, et al.
Pubblicazione: (2025)
di: Liu, Jincheng, et al.
Pubblicazione: (2025)
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
di: Kouremetis, Michael, et al.
Pubblicazione: (2025)
di: Kouremetis, Michael, et al.
Pubblicazione: (2025)
LARP: Language-Agent Role Play for Open-World Games
di: Yan, Ming, et al.
Pubblicazione: (2023)
di: Yan, Ming, et al.
Pubblicazione: (2023)
MINDECHO: Role-Playing Language Agents for Key Opinion Leaders
di: Xu, Rui, et al.
Pubblicazione: (2024)
di: Xu, Rui, et al.
Pubblicazione: (2024)
Knowledge Graph-enhanced Large Language Model for Incremental Game PlayTesting
di: Mu, Enhong, et al.
Pubblicazione: (2025)
di: Mu, Enhong, et al.
Pubblicazione: (2025)
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
di: Zhang, Xiaotian, et al.
Pubblicazione: (2023)
di: Zhang, Xiaotian, et al.
Pubblicazione: (2023)
Open-Set Living Need Prediction with Large Language Models
di: Lan, Xiaochong, et al.
Pubblicazione: (2025)
di: Lan, Xiaochong, et al.
Pubblicazione: (2025)
Evaluating Reliability Gaps in Large Language Model Safety via Repeated Prompt Sampling
di: Broadwater, Keita
Pubblicazione: (2026)
di: Broadwater, Keita
Pubblicazione: (2026)
LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social Simulations
di: Pham, Viet-Thanh, et al.
Pubblicazione: (2026)
di: Pham, Viet-Thanh, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Large Language Models as Pokémon Battle Agents: Strategic Play and Content Generation
di: Jain, Daksh, et al.
Pubblicazione: (2025) -
A Multi-Agent Pokemon Tournament for Evaluating Strategic Reasoning of Large Language Models
di: Yashwanth, Tadisetty Sai, et al.
Pubblicazione: (2025) -
Role-Playing Evaluation for Large Language Models
di: Boudouri, Yassine El, et al.
Pubblicazione: (2025) -
Language Agents with Reinforcement Learning for Strategic Play in the Werewolf Game
di: Xu, Zelai, et al.
Pubblicazione: (2023) -
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
di: Wang, Kevin, et al.
Pubblicazione: (2026)