Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Light, Jonathan, Cai, Min, Chen, Weiqin, Wang, Guanzhi, Chen, Xiusi, Cheng, Wei, Yue, Yisong, Hu, Ziniu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908469716582400
author Light, Jonathan
Cai, Min
Chen, Weiqin
Wang, Guanzhi
Chen, Xiusi
Cheng, Wei
Yue, Yisong
Hu, Ziniu
author_facet Light, Jonathan
Cai, Min
Chen, Weiqin
Wang, Guanzhi
Chen, Xiusi
Cheng, Wei
Yue, Yisong
Hu, Ziniu
contents Traditional reinforcement learning and planning typically requires vast amounts of data and training to develop effective policies. In contrast, large language models (LLMs) exhibit strong generalization and zero-shot capabilities, but struggle with tasks that require detailed planning and decision-making in complex action spaces. We introduce STRATEGIST, a novel approach that integrates the strengths of both methods. Our approach leverages LLMs to search and update high-level strategies (as text), which are then refined and executed by low-level Monte Carlo Tree Search (MCTS). STRATEGIST is a generalizable framework to optimize the strategy through population-based self-play simulations without the need for any training data. We demonstrate the effectiveness of STRATEGIST in learning optimal strategies for competitive, multi-turn games with partial information, including Game of Pure Strategy (GOPS) and multi-agent, hidden-identity discussion games like The Resistance: Avalon. Our results show that agents equipped with STRATEGIST outperform those trained with traditional RL methods, other LLM-based skill acquisition techniques, pre-existing LLM agents across both game environments and achieves comparable performance against human players.
format Preprint
id arxiv_https___arxiv_org_abs_2408_10635
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search
Light, Jonathan
Cai, Min
Chen, Weiqin
Wang, Guanzhi
Chen, Xiusi
Cheng, Wei
Yue, Yisong
Hu, Ziniu
Artificial Intelligence
Computation and Language
Traditional reinforcement learning and planning typically requires vast amounts of data and training to develop effective policies. In contrast, large language models (LLMs) exhibit strong generalization and zero-shot capabilities, but struggle with tasks that require detailed planning and decision-making in complex action spaces. We introduce STRATEGIST, a novel approach that integrates the strengths of both methods. Our approach leverages LLMs to search and update high-level strategies (as text), which are then refined and executed by low-level Monte Carlo Tree Search (MCTS). STRATEGIST is a generalizable framework to optimize the strategy through population-based self-play simulations without the need for any training data. We demonstrate the effectiveness of STRATEGIST in learning optimal strategies for competitive, multi-turn games with partial information, including Game of Pure Strategy (GOPS) and multi-agent, hidden-identity discussion games like The Resistance: Avalon. Our results show that agents equipped with STRATEGIST outperform those trained with traditional RL methods, other LLM-based skill acquisition techniques, pre-existing LLM agents across both game environments and achieves comparable performance against human players.
title Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2408.10635