Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908469716582400 |
|---|---|
| author | Light, Jonathan Cai, Min Chen, Weiqin Wang, Guanzhi Chen, Xiusi Cheng, Wei Yue, Yisong Hu, Ziniu |
| author_facet | Light, Jonathan Cai, Min Chen, Weiqin Wang, Guanzhi Chen, Xiusi Cheng, Wei Yue, Yisong Hu, Ziniu |
| contents | Traditional reinforcement learning and planning typically requires vast amounts of data and training to develop effective policies. In contrast, large language models (LLMs) exhibit strong generalization and zero-shot capabilities, but struggle with tasks that require detailed planning and decision-making in complex action spaces. We introduce STRATEGIST, a novel approach that integrates the strengths of both methods. Our approach leverages LLMs to search and update high-level strategies (as text), which are then refined and executed by low-level Monte Carlo Tree Search (MCTS). STRATEGIST is a generalizable framework to optimize the strategy through population-based self-play simulations without the need for any training data. We demonstrate the effectiveness of STRATEGIST in learning optimal strategies for competitive, multi-turn games with partial information, including Game of Pure Strategy (GOPS) and multi-agent, hidden-identity discussion games like The Resistance: Avalon. Our results show that agents equipped with STRATEGIST outperform those trained with traditional RL methods, other LLM-based skill acquisition techniques, pre-existing LLM agents across both game environments and achieves comparable performance against human players. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_10635 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search Light, Jonathan Cai, Min Chen, Weiqin Wang, Guanzhi Chen, Xiusi Cheng, Wei Yue, Yisong Hu, Ziniu Artificial Intelligence Computation and Language Traditional reinforcement learning and planning typically requires vast amounts of data and training to develop effective policies. In contrast, large language models (LLMs) exhibit strong generalization and zero-shot capabilities, but struggle with tasks that require detailed planning and decision-making in complex action spaces. We introduce STRATEGIST, a novel approach that integrates the strengths of both methods. Our approach leverages LLMs to search and update high-level strategies (as text), which are then refined and executed by low-level Monte Carlo Tree Search (MCTS). STRATEGIST is a generalizable framework to optimize the strategy through population-based self-play simulations without the need for any training data. We demonstrate the effectiveness of STRATEGIST in learning optimal strategies for competitive, multi-turn games with partial information, including Game of Pure Strategy (GOPS) and multi-agent, hidden-identity discussion games like The Resistance: Avalon. Our results show that agents equipped with STRATEGIST outperform those trained with traditional RL methods, other LLM-based skill acquisition techniques, pre-existing LLM agents across both game environments and achieves comparable performance against human players. |
| title | Strategist: Self-improvement of LLM Decision Making via Bi-Level Tree Search |
| topic | Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2408.10635 |