Look Further Ahead: Testing the Limits of GPT-4 in Path Planning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aghzal, Mohamed, Plaku, Erion, Yao, Ziyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914843262451712
author Aghzal, Mohamed
Plaku, Erion
Yao, Ziyu
author_facet Aghzal, Mohamed
Plaku, Erion
Yao, Ziyu
contents Large Language Models (LLMs) have shown impressive capabilities across a wide variety of tasks. However, they still face challenges with long-horizon planning. To study this, we propose path planning tasks as a platform to evaluate LLMs' ability to navigate long trajectories under geometric constraints. Our proposed benchmark systematically tests path-planning skills in complex settings. Using this, we examined GPT-4's planning abilities using various task representations and prompting approaches. We found that framing prompts as Python code and decomposing long trajectory tasks improve GPT-4's path planning effectiveness. However, while these approaches show some promise toward improving the planning ability of the model, they do not obtain optimal paths and fail at generalizing over extended horizons.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12000
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Look Further Ahead: Testing the Limits of GPT-4 in Path Planning
Aghzal, Mohamed
Plaku, Erion
Yao, Ziyu
Artificial Intelligence
Large Language Models (LLMs) have shown impressive capabilities across a wide variety of tasks. However, they still face challenges with long-horizon planning. To study this, we propose path planning tasks as a platform to evaluate LLMs' ability to navigate long trajectories under geometric constraints. Our proposed benchmark systematically tests path-planning skills in complex settings. Using this, we examined GPT-4's planning abilities using various task representations and prompting approaches. We found that framing prompts as Python code and decomposing long trajectory tasks improve GPT-4's path planning effectiveness. However, while these approaches show some promise toward improving the planning ability of the model, they do not obtain optimal paths and fail at generalizing over extended horizons.
title Look Further Ahead: Testing the Limits of GPT-4 in Path Planning
topic Artificial Intelligence
url https://arxiv.org/abs/2406.12000