Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Weiyu, Mi, Qirui, Zeng, Yongcheng, Yan, Xue, Wu, Yuqiao, Lin, Runji, Zhang, Haifeng, Wang, Jun
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914837157642240
author Ma, Weiyu
Mi, Qirui
Zeng, Yongcheng
Yan, Xue
Wu, Yuqiao
Lin, Runji
Zhang, Haifeng
Wang, Jun
author_facet Ma, Weiyu
Mi, Qirui
Zeng, Yongcheng
Yan, Xue
Wu, Yuqiao
Lin, Runji
Zhang, Haifeng
Wang, Jun
contents StarCraft II is a challenging benchmark for AI agents due to the necessity of both precise micro level operations and strategic macro awareness. Previous works, such as Alphastar and SCC, achieve impressive performance on tackling StarCraft II , however, still exhibit deficiencies in long term strategic planning and strategy interpretability. Emerging large language model (LLM) agents, such as Voyage and MetaGPT, presents the immense potential in solving intricate tasks. Motivated by this, we aim to validate the capabilities of LLMs on StarCraft II, a highly complex RTS game.To conveniently take full advantage of LLMs` reasoning abilities, we first develop textual StratCraft II environment, called TextStarCraft II, which LLM agent can interact. Secondly, we propose a Chain of Summarization method, including single frame summarization for processing raw observations and multi frame summarization for analyzing game information, providing command recommendations, and generating strategic decisions. Our experiment consists of two parts: first, an evaluation by human experts, which includes assessing the LLMs`s mastery of StarCraft II knowledge and the performance of LLM agents in the game; second, the in game performance of LLM agents, encompassing aspects like win rate and the impact of Chain of Summarization.Experiment results demonstrate that: 1. LLMs possess the relevant knowledge and complex planning abilities needed to address StarCraft II scenarios; 2. Human experts consider the performance of LLM agents to be close to that of an average player who has played StarCraft II for eight years; 3. LLM agents are capable of defeating the built in AI at the Harder(Lv5) difficulty level. We have open sourced the code and released demo videos of LLM agent playing StarCraft II.
format Preprint
id arxiv_https___arxiv_org_abs_2312_11865
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach
Ma, Weiyu
Mi, Qirui
Zeng, Yongcheng
Yan, Xue
Wu, Yuqiao
Lin, Runji
Zhang, Haifeng
Wang, Jun
Artificial Intelligence
StarCraft II is a challenging benchmark for AI agents due to the necessity of both precise micro level operations and strategic macro awareness. Previous works, such as Alphastar and SCC, achieve impressive performance on tackling StarCraft II , however, still exhibit deficiencies in long term strategic planning and strategy interpretability. Emerging large language model (LLM) agents, such as Voyage and MetaGPT, presents the immense potential in solving intricate tasks. Motivated by this, we aim to validate the capabilities of LLMs on StarCraft II, a highly complex RTS game.To conveniently take full advantage of LLMs` reasoning abilities, we first develop textual StratCraft II environment, called TextStarCraft II, which LLM agent can interact. Secondly, we propose a Chain of Summarization method, including single frame summarization for processing raw observations and multi frame summarization for analyzing game information, providing command recommendations, and generating strategic decisions. Our experiment consists of two parts: first, an evaluation by human experts, which includes assessing the LLMs`s mastery of StarCraft II knowledge and the performance of LLM agents in the game; second, the in game performance of LLM agents, encompassing aspects like win rate and the impact of Chain of Summarization.Experiment results demonstrate that: 1. LLMs possess the relevant knowledge and complex planning abilities needed to address StarCraft II scenarios; 2. Human experts consider the performance of LLM agents to be close to that of an average player who has played StarCraft II for eight years; 3. LLM agents are capable of defeating the built in AI at the Harder(Lv5) difficulty level. We have open sourced the code and released demo videos of LLM agent playing StarCraft II.
title Large Language Models Play StarCraft II: Benchmarks and A Chain of Summarization Approach
topic Artificial Intelligence
url https://arxiv.org/abs/2312.11865