Planetarium: A Rigorous Benchmark for Translating Text to Structured Planning Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zuo, Max, Velez, Francisco Piedrahita, Li, Xiaochen, Littman, Michael L., Bach, Stephen H.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915610297892864
author Zuo, Max
Velez, Francisco Piedrahita
Li, Xiaochen
Littman, Michael L.
Bach, Stephen H.
author_facet Zuo, Max
Velez, Francisco Piedrahita
Li, Xiaochen
Littman, Michael L.
Bach, Stephen H.
contents Recent works have explored using language models for planning problems. One approach examines translating natural language descriptions of planning tasks into structured planning languages, such as the planning domain definition language (PDDL). Existing evaluation methods struggle to ensure semantic correctness and rely on simple or unrealistic datasets. To bridge this gap, we introduce \textit{Planetarium}, a benchmark designed to evaluate language models' ability to generate PDDL code from natural language descriptions of planning tasks. \textit{Planetarium} features a novel PDDL equivalence algorithm that flexibly evaluates the correctness of generated PDDL, along with a dataset of 145,918 text-to-PDDL pairs across 73 unique state combinations with varying levels of difficulty. Finally, we evaluate several API-access and open-weight language models that reveal this task's complexity. For example, 96.1\% of the PDDL problem descriptions generated by GPT-4o are syntactically parseable, 94.4\% are solvable, but only 24.8\% are semantically correct, highlighting the need for a more rigorous benchmark for this problem.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03321
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Planetarium: A Rigorous Benchmark for Translating Text to Structured Planning Languages
Zuo, Max
Velez, Francisco Piedrahita
Li, Xiaochen
Littman, Michael L.
Bach, Stephen H.
Computation and Language
Artificial Intelligence
Machine Learning
Recent works have explored using language models for planning problems. One approach examines translating natural language descriptions of planning tasks into structured planning languages, such as the planning domain definition language (PDDL). Existing evaluation methods struggle to ensure semantic correctness and rely on simple or unrealistic datasets. To bridge this gap, we introduce \textit{Planetarium}, a benchmark designed to evaluate language models' ability to generate PDDL code from natural language descriptions of planning tasks. \textit{Planetarium} features a novel PDDL equivalence algorithm that flexibly evaluates the correctness of generated PDDL, along with a dataset of 145,918 text-to-PDDL pairs across 73 unique state combinations with varying levels of difficulty. Finally, we evaluate several API-access and open-weight language models that reveal this task's complexity. For example, 96.1\% of the PDDL problem descriptions generated by GPT-4o are syntactically parseable, 94.4\% are solvable, but only 24.8\% are semantically correct, highlighting the need for a more rigorous benchmark for this problem.
title Planetarium: A Rigorous Benchmark for Translating Text to Structured Planning Languages
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.03321