Saved in:
Bibliographic Details
Main Authors: Qu, Yun, Wang, Boyuan, Jiang, Yuhang, Shao, Jianzhun, Mao, Yixiu, Zou, Heming, Liu, Chang, Wang, Cheems, Liu, Meiqin, Ji, Xiangyang
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2410.02511
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914619820343296
author Qu, Yun
Wang, Boyuan
Jiang, Yuhang
Shao, Jianzhun
Mao, Yixiu
Zou, Heming
Liu, Chang
Wang, Cheems
Liu, Meiqin
Ji, Xiangyang
author_facet Qu, Yun
Wang, Boyuan
Jiang, Yuhang
Shao, Jianzhun
Mao, Yixiu
Zou, Heming
Liu, Chang
Wang, Cheems
Liu, Meiqin
Ji, Xiangyang
contents With expansive state-action spaces, efficient multi-agent exploration remains a longstanding challenge in reinforcement learning. Although pursuing novelty, diversity, or uncertainty attracts increasing attention, redundant efforts brought by exploration without proper guidance choices poses a practical issue for the community. This paper introduces a systematic approach, termed LEMAE, choosing to channel informative task-relevant guidance from a knowledgeable Large Language Model (LLM) for Efficient Multi-Agent Exploration. Specifically, we ground linguistic knowledge from LLM into symbolic key states, that are critical for task fulfillment, in a discriminative manner at low LLM inference costs. To unleash the power of key states, we design Subspace-based Hindsight Intrinsic Reward (SHIR) to guide agents toward key states by increasing reward density. Additionally, we build the Key State Memory Tree (KSMT) to track transitions between key states in a specific task for organized exploration. Benefiting from diminishing redundant explorations, LEMAE outperforms existing SOTA approaches on the challenging benchmarks (e.g., SMAC and MPE) by a large margin, achieving a 10x acceleration in certain scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2410_02511
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Stop Wandering, Find the Keys: LLMs Discriminate Key States for Efficient Multi-Agent Exploration
Qu, Yun
Wang, Boyuan
Jiang, Yuhang
Shao, Jianzhun
Mao, Yixiu
Zou, Heming
Liu, Chang
Wang, Cheems
Liu, Meiqin
Ji, Xiangyang
Artificial Intelligence
Multiagent Systems
With expansive state-action spaces, efficient multi-agent exploration remains a longstanding challenge in reinforcement learning. Although pursuing novelty, diversity, or uncertainty attracts increasing attention, redundant efforts brought by exploration without proper guidance choices poses a practical issue for the community. This paper introduces a systematic approach, termed LEMAE, choosing to channel informative task-relevant guidance from a knowledgeable Large Language Model (LLM) for Efficient Multi-Agent Exploration. Specifically, we ground linguistic knowledge from LLM into symbolic key states, that are critical for task fulfillment, in a discriminative manner at low LLM inference costs. To unleash the power of key states, we design Subspace-based Hindsight Intrinsic Reward (SHIR) to guide agents toward key states by increasing reward density. Additionally, we build the Key State Memory Tree (KSMT) to track transitions between key states in a specific task for organized exploration. Benefiting from diminishing redundant explorations, LEMAE outperforms existing SOTA approaches on the challenging benchmarks (e.g., SMAC and MPE) by a large margin, achieving a 10x acceleration in certain scenarios.
title Stop Wandering, Find the Keys: LLMs Discriminate Key States for Efficient Multi-Agent Exploration
topic Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2410.02511