PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908524509921280 |
|---|---|
| author | Schipper, Olivier Zhang, Yudi Du, Yali Pechenizkiy, Mykola Fang, Meng |
| author_facet | Schipper, Olivier Zhang, Yudi Du, Yali Pechenizkiy, Mykola Fang, Meng |
| contents | LLM-based agents have shown promise in various cooperative and strategic reasoning tasks, but their effectiveness in competitive multi-agent environments remains underexplored. To address this gap, we introduce PillagerBench, a novel framework for evaluating multi-agent systems in real-time competitive team-vs-team scenarios in Minecraft. It provides an extensible API, multi-round testing, and rule-based built-in opponents for fair, reproducible comparisons. We also propose TactiCrafter, an LLM-based multi-agent system that facilitates teamwork through human-readable tactics, learns causal dependencies, and adapts to opponent strategies. Our evaluation demonstrates that TactiCrafter outperforms baseline approaches and showcases adaptive learning through self-play. Additionally, we analyze its learning process and strategic evolution over multiple game episodes. To encourage further research, we have open-sourced PillagerBench, fostering advancements in multi-agent AI for competitive environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_06235 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments Schipper, Olivier Zhang, Yudi Du, Yali Pechenizkiy, Mykola Fang, Meng Artificial Intelligence Multiagent Systems I.2.11; I.2.6; I.2.8 LLM-based agents have shown promise in various cooperative and strategic reasoning tasks, but their effectiveness in competitive multi-agent environments remains underexplored. To address this gap, we introduce PillagerBench, a novel framework for evaluating multi-agent systems in real-time competitive team-vs-team scenarios in Minecraft. It provides an extensible API, multi-round testing, and rule-based built-in opponents for fair, reproducible comparisons. We also propose TactiCrafter, an LLM-based multi-agent system that facilitates teamwork through human-readable tactics, learns causal dependencies, and adapts to opponent strategies. Our evaluation demonstrates that TactiCrafter outperforms baseline approaches and showcases adaptive learning through self-play. Additionally, we analyze its learning process and strategic evolution over multiple game episodes. To encourage further research, we have open-sourced PillagerBench, fostering advancements in multi-agent AI for competitive environments. |
| title | PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments |
| topic | Artificial Intelligence Multiagent Systems I.2.11; I.2.6; I.2.8 |
| url | https://arxiv.org/abs/2509.06235 |