MCU: An Evaluation Framework for Open-Ended Game Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zheng, Xinyue, Lin, Haowei, He, Kaichen, Wang, Zihao, Zheng, Zilong, Liang, Yitao
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912409751388160
author Zheng, Xinyue
Lin, Haowei
He, Kaichen
Wang, Zihao
Zheng, Zilong
Liang, Yitao
author_facet Zheng, Xinyue
Lin, Haowei
He, Kaichen
Wang, Zihao
Zheng, Zilong
Liang, Yitao
contents Developing AI agents capable of interacting with open-world environments to solve diverse tasks is a compelling challenge. However, evaluating such open-ended agents remains difficult, with current benchmarks facing scalability limitations. To address this, we introduce Minecraft Universe (MCU), a comprehensive evaluation framework set within the open-world video game Minecraft. MCU incorporates three key components: (1) an expanding collection of 3,452 composable atomic tasks that encompasses 11 major categories and 41 subcategories of challenges; (2) a task composition mechanism capable of generating infinite diverse tasks with varying difficulty; and (3) a general evaluation framework that achieves 91.5\% alignment with human ratings for open-ended task assessment. Empirical results reveal that even state-of-the-art foundation agents struggle with the increasing diversity and complexity of tasks. These findings highlight the necessity of MCU as a robust benchmark to drive progress in AI agent development within open-ended environments. Our evaluation code and scripts are available at https://github.com/CraftJarvis/MCU.
format Preprint
id arxiv_https___arxiv_org_abs_2310_08367
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MCU: An Evaluation Framework for Open-Ended Game Agents
Zheng, Xinyue
Lin, Haowei
He, Kaichen
Wang, Zihao
Zheng, Zilong
Liang, Yitao
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
Developing AI agents capable of interacting with open-world environments to solve diverse tasks is a compelling challenge. However, evaluating such open-ended agents remains difficult, with current benchmarks facing scalability limitations. To address this, we introduce Minecraft Universe (MCU), a comprehensive evaluation framework set within the open-world video game Minecraft. MCU incorporates three key components: (1) an expanding collection of 3,452 composable atomic tasks that encompasses 11 major categories and 41 subcategories of challenges; (2) a task composition mechanism capable of generating infinite diverse tasks with varying difficulty; and (3) a general evaluation framework that achieves 91.5\% alignment with human ratings for open-ended task assessment. Empirical results reveal that even state-of-the-art foundation agents struggle with the increasing diversity and complexity of tasks. These findings highlight the necessity of MCU as a robust benchmark to drive progress in AI agent development within open-ended environments. Our evaluation code and scripts are available at https://github.com/CraftJarvis/MCU.
title MCU: An Evaluation Framework for Open-Ended Game Agents
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2310.08367