Exploring and Benchmarking the Planning Capabilities of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bohnet, Bernd, Nova, Azade, Parisi, Aaron T, Swersky, Kevin, Goshvadi, Katayoon, Dai, Hanjun, Schuurmans, Dale, Fiedel, Noah, Sedghi, Hanie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929573324652544
author Bohnet, Bernd
Nova, Azade
Parisi, Aaron T
Swersky, Kevin
Goshvadi, Katayoon
Dai, Hanjun
Schuurmans, Dale
Fiedel, Noah
Sedghi, Hanie
author_facet Bohnet, Bernd
Nova, Azade
Parisi, Aaron T
Swersky, Kevin
Goshvadi, Katayoon
Dai, Hanjun
Schuurmans, Dale
Fiedel, Noah
Sedghi, Hanie
contents Classical and natural language planning tasks remain a difficult domain for modern large language models (LLMs). In this work, we lay the foundations for improving planning capabilities of LLMs. First, we construct a comprehensive benchmark suite encompassing both classical planning benchmarks and natural language scenarios. This suite includes algorithms to methodically generate instances of tasks with varying levels of difficulty, allowing for rigorous and systematic evaluation of LLM performance. Next, we investigate the use of many-shot in-context learning to enhance LLM planning, exploring the relationship between increased context length and improved planning performance. In addition, we demonstrate the positive impact of fine-tuning LLMs on optimal planning paths. We also probe the efficacy of chain-of-thought reasoning methods to improve LLM planning performance. Moreover, we probe the performance of the proposed methods in out-of-distribution scenarios, assessing the ability to generalize to novel and unseen planning challenges. Finally, we investigate model's failure modes and reveal insights that hold true across different benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13094
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring and Benchmarking the Planning Capabilities of Large Language Models
Bohnet, Bernd
Nova, Azade
Parisi, Aaron T
Swersky, Kevin
Goshvadi, Katayoon
Dai, Hanjun
Schuurmans, Dale
Fiedel, Noah
Sedghi, Hanie
Computation and Language
Artificial Intelligence
Machine Learning
Classical and natural language planning tasks remain a difficult domain for modern large language models (LLMs). In this work, we lay the foundations for improving planning capabilities of LLMs. First, we construct a comprehensive benchmark suite encompassing both classical planning benchmarks and natural language scenarios. This suite includes algorithms to methodically generate instances of tasks with varying levels of difficulty, allowing for rigorous and systematic evaluation of LLM performance. Next, we investigate the use of many-shot in-context learning to enhance LLM planning, exploring the relationship between increased context length and improved planning performance. In addition, we demonstrate the positive impact of fine-tuning LLMs on optimal planning paths. We also probe the efficacy of chain-of-thought reasoning methods to improve LLM planning performance. Moreover, we probe the performance of the proposed methods in out-of-distribution scenarios, assessing the ability to generalize to novel and unseen planning challenges. Finally, we investigate model's failure modes and reveal insights that hold true across different benchmarks.
title Exploring and Benchmarking the Planning Capabilities of Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.13094