PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schipper, Olivier, Zhang, Yudi, Du, Yali, Pechenizkiy, Mykola, Fang, Meng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908524509921280
author Schipper, Olivier
Zhang, Yudi
Du, Yali
Pechenizkiy, Mykola
Fang, Meng
author_facet Schipper, Olivier
Zhang, Yudi
Du, Yali
Pechenizkiy, Mykola
Fang, Meng
contents LLM-based agents have shown promise in various cooperative and strategic reasoning tasks, but their effectiveness in competitive multi-agent environments remains underexplored. To address this gap, we introduce PillagerBench, a novel framework for evaluating multi-agent systems in real-time competitive team-vs-team scenarios in Minecraft. It provides an extensible API, multi-round testing, and rule-based built-in opponents for fair, reproducible comparisons. We also propose TactiCrafter, an LLM-based multi-agent system that facilitates teamwork through human-readable tactics, learns causal dependencies, and adapts to opponent strategies. Our evaluation demonstrates that TactiCrafter outperforms baseline approaches and showcases adaptive learning through self-play. Additionally, we analyze its learning process and strategic evolution over multiple game episodes. To encourage further research, we have open-sourced PillagerBench, fostering advancements in multi-agent AI for competitive environments.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06235
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments
Schipper, Olivier
Zhang, Yudi
Du, Yali
Pechenizkiy, Mykola
Fang, Meng
Artificial Intelligence
Multiagent Systems
I.2.11; I.2.6; I.2.8
LLM-based agents have shown promise in various cooperative and strategic reasoning tasks, but their effectiveness in competitive multi-agent environments remains underexplored. To address this gap, we introduce PillagerBench, a novel framework for evaluating multi-agent systems in real-time competitive team-vs-team scenarios in Minecraft. It provides an extensible API, multi-round testing, and rule-based built-in opponents for fair, reproducible comparisons. We also propose TactiCrafter, an LLM-based multi-agent system that facilitates teamwork through human-readable tactics, learns causal dependencies, and adapts to opponent strategies. Our evaluation demonstrates that TactiCrafter outperforms baseline approaches and showcases adaptive learning through self-play. Additionally, we analyze its learning process and strategic evolution over multiple game episodes. To encourage further research, we have open-sourced PillagerBench, fostering advancements in multi-agent AI for competitive environments.
title PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments
topic Artificial Intelligence
Multiagent Systems
I.2.11; I.2.6; I.2.8
url https://arxiv.org/abs/2509.06235