A Multi-Agent Pokemon Tournament for Evaluating Strategic Reasoning of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yashwanth, Tadisetty Sai, C, Dhatri |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Structure of Floating-Point Noise in Batch-Invariant GPU Matrix Multiplication
von: Yashwanth, Tadisetty Sai
Veröffentlicht: (2025)
von: Yashwanth, Tadisetty Sai
Veröffentlicht: (2025)
Large Language Models as Pokémon Battle Agents: Strategic Play and Content Generation
von: Jain, Daksh, et al.
Veröffentlicht: (2025)
von: Jain, Daksh, et al.
Veröffentlicht: (2025)
Real Time Child Abduction And Detection System
von: Yashwanth, Tadisetty Sai, et al.
Veröffentlicht: (2025)
von: Yashwanth, Tadisetty Sai, et al.
Veröffentlicht: (2025)
PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models
von: Hu, Sihao, et al.
Veröffentlicht: (2024)
von: Hu, Sihao, et al.
Veröffentlicht: (2024)
The Aegis Protocol: A Foundational Security Framework for Autonomous AI Agents
von: Adapala, Sai Teja Reddy, et al.
Veröffentlicht: (2025)
von: Adapala, Sai Teja Reddy, et al.
Veröffentlicht: (2025)
Evaluating Strategic Reasoning in Forecasting Agents
von: Liptay, Tom, et al.
Veröffentlicht: (2026)
von: Liptay, Tom, et al.
Veröffentlicht: (2026)
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
von: Wang, Kevin, et al.
Veröffentlicht: (2026)
von: Wang, Kevin, et al.
Veröffentlicht: (2026)
Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games
von: He, Yidong, et al.
Veröffentlicht: (2026)
von: He, Yidong, et al.
Veröffentlicht: (2026)
Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play
von: Ekne, H. C.
Veröffentlicht: (2026)
von: Ekne, H. C.
Veröffentlicht: (2026)
ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models
von: Liu, Jincheng, et al.
Veröffentlicht: (2025)
von: Liu, Jincheng, et al.
Veröffentlicht: (2025)
PTCG-Bench: Can LLM Agents Master Pokémon Trading Card Game?
von: Hua, Dongdong, et al.
Veröffentlicht: (2026)
von: Hua, Dongdong, et al.
Veröffentlicht: (2026)
SV3.3B: A Sports Video Understanding Model for Action Recognition
von: Kodathala, Sai Varun, et al.
Veröffentlicht: (2025)
von: Kodathala, Sai Varun, et al.
Veröffentlicht: (2025)
Arena-Lite: Efficient and Reliable Large Language Model Evaluation via Tournament-Based Direct Comparisons
von: Son, Seonil, et al.
Veröffentlicht: (2024)
von: Son, Seonil, et al.
Veröffentlicht: (2024)
Responsibility-aware Strategic Reasoning in Probabilistic Multi-Agent Systems
von: Mu, Chunyan, et al.
Veröffentlicht: (2024)
von: Mu, Chunyan, et al.
Veröffentlicht: (2024)
CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
von: Costarelli, Anthony, et al.
Veröffentlicht: (2024)
von: Costarelli, Anthony, et al.
Veröffentlicht: (2024)
Cooperative Strategic Planning Enhances Reasoning Capabilities in Large Language Models
von: Wang, Danqing, et al.
Veröffentlicht: (2024)
von: Wang, Danqing, et al.
Veröffentlicht: (2024)
Crimson: Empowering Strategic Reasoning in Cybersecurity through Large Language Models
von: Jin, Jiandong, et al.
Veröffentlicht: (2024)
von: Jin, Jiandong, et al.
Veröffentlicht: (2024)
MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
von: Yuan, Huining, et al.
Veröffentlicht: (2025)
von: Yuan, Huining, et al.
Veröffentlicht: (2025)
Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models
von: Wit, Maria Carolina Cornelia, et al.
Veröffentlicht: (2025)
von: Wit, Maria Carolina Cornelia, et al.
Veröffentlicht: (2025)
K-Level Reasoning: Establishing Higher Order Beliefs in Large Language Models for Strategic Reasoning
von: Zhang, Yadong, et al.
Veröffentlicht: (2024)
von: Zhang, Yadong, et al.
Veröffentlicht: (2024)
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models
von: Yin, Qiyue, et al.
Veröffentlicht: (2025)
von: Yin, Qiyue, et al.
Veröffentlicht: (2025)
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
von: Xu, Zelai, et al.
Veröffentlicht: (2025)
von: Xu, Zelai, et al.
Veröffentlicht: (2025)
Neural Algorithmic Reasoners informed Large Language Model for Multi-Agent Path Finding
von: Feng, Pu, et al.
Veröffentlicht: (2025)
von: Feng, Pu, et al.
Veröffentlicht: (2025)
Integrating Graphs, Large Language Models, and Agents: Reasoning and Retrieval
von: Jelodar, Hamed, et al.
Veröffentlicht: (2026)
von: Jelodar, Hamed, et al.
Veröffentlicht: (2026)
Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
von: Adapala, Sai Teja Reddy
Veröffentlicht: (2025)
von: Adapala, Sai Teja Reddy
Veröffentlicht: (2025)
MedLA: A Logic-Driven Multi-Agent Framework for Complex Medical Reasoning with Large Language Models
von: Ma, Siqi, et al.
Veröffentlicht: (2025)
von: Ma, Siqi, et al.
Veröffentlicht: (2025)
DixitWorld: Evaluating Multimodal Abductive Reasoning in Vision-Language Models with Multi-Agent Dixit Gameplay
von: Mo, Yunxiang, et al.
Veröffentlicht: (2025)
von: Mo, Yunxiang, et al.
Veröffentlicht: (2025)
Tree of Agents: Improving Long-Context Capabilities of Large Language Models through Multi-Perspective Reasoning
von: Yu, Song, et al.
Veröffentlicht: (2025)
von: Yu, Song, et al.
Veröffentlicht: (2025)
Society of Mind Meets Real-Time Strategy: A Hierarchical Multi-Agent Framework for Strategic Reasoning
von: Ahn, Daechul, et al.
Veröffentlicht: (2025)
von: Ahn, Daechul, et al.
Veröffentlicht: (2025)
PokéAI: A Goal-Generating, Battle-Optimizing Multi-agent System for Pokemon Red
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
von: Lu, Leo, et al.
Veröffentlicht: (2025)
von: Lu, Leo, et al.
Veröffentlicht: (2025)
METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models
von: Li, Pengfeng, et al.
Veröffentlicht: (2026)
von: Li, Pengfeng, et al.
Veröffentlicht: (2026)
JudgeSQL: Reasoning over SQL Candidates with Weighted Consensus Tournament
von: Bai, Jiayuan, et al.
Veröffentlicht: (2025)
von: Bai, Jiayuan, et al.
Veröffentlicht: (2025)
LLM-Guided Synthetic Augmentation (LGSA) for Mitigating Bias in AI Systems
von: Karri, Sai Suhruth Reddy, et al.
Veröffentlicht: (2025)
von: Karri, Sai Suhruth Reddy, et al.
Veröffentlicht: (2025)
Towards Collaborative Intelligence: Propagating Intentions and Reasoning for Multi-Agent Coordination with Large Language Models
von: Qiu, Xihe, et al.
Veröffentlicht: (2024)
von: Qiu, Xihe, et al.
Veröffentlicht: (2024)
An Automated Multi-modal Evaluation Framework for Mobile Intelligent Assistants Based on Large Language Models and Multi-Agent Collaboration
von: Wang, Meiping, et al.
Veröffentlicht: (2025)
von: Wang, Meiping, et al.
Veröffentlicht: (2025)
A Comprehensive Evaluation on Event Reasoning of Large Language Models
von: Tao, Zhengwei, et al.
Veröffentlicht: (2024)
von: Tao, Zhengwei, et al.
Veröffentlicht: (2024)
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
von: Patel, Nisarg, et al.
Veröffentlicht: (2024)
von: Patel, Nisarg, et al.
Veröffentlicht: (2024)
Improving Physics Reasoning in Large Language Models Using Mixture of Refinement Agents
von: Jaiswal, Raj, et al.
Veröffentlicht: (2024)
von: Jaiswal, Raj, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On the Structure of Floating-Point Noise in Batch-Invariant GPU Matrix Multiplication
von: Yashwanth, Tadisetty Sai
Veröffentlicht: (2025) -
Large Language Models as Pokémon Battle Agents: Strategic Play and Content Generation
von: Jain, Daksh, et al.
Veröffentlicht: (2025) -
Real Time Child Abduction And Detection System
von: Yashwanth, Tadisetty Sai, et al.
Veröffentlicht: (2025) -
PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models
von: Hu, Sihao, et al.
Veröffentlicht: (2024) -
The Aegis Protocol: A Foundational Security Framework for Autonomous AI Agents
von: Adapala, Sai Teja Reddy, et al.
Veröffentlicht: (2025)