An Extensive Evaluation of PDDL Capabilities in off-the-shelf LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Vyas, Kaustubh, Graux, Damien, Montella, Sébastien, Vougiouklis, Pavlos, Lai, Ruofei, Li, Keshuang, Ren, Yang, Pan, Jeff Z.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915175749124096
author Vyas, Kaustubh
Graux, Damien
Montella, Sébastien
Vougiouklis, Pavlos
Lai, Ruofei
Li, Keshuang
Ren, Yang
Pan, Jeff Z.
author_facet Vyas, Kaustubh
Graux, Damien
Montella, Sébastien
Vougiouklis, Pavlos
Lai, Ruofei
Li, Keshuang
Ren, Yang
Pan, Jeff Z.
contents In recent advancements, large language models (LLMs) have exhibited proficiency in code generation and chain-of-thought reasoning, laying the groundwork for tackling automatic formal planning tasks. This study evaluates the potential of LLMs to understand and generate Planning Domain Definition Language (PDDL), an essential representation in artificial intelligence planning. We conduct an extensive analysis across 20 distinct models spanning 7 major LLM families, both commercial and open-source. Our comprehensive evaluation sheds light on the zero-shot LLM capabilities of parsing, generating, and reasoning with PDDL. Our findings indicate that while some models demonstrate notable effectiveness in handling PDDL, others pose limitations in more complex scenarios requiring nuanced planning knowledge. These results highlight the promise and current limitations of LLMs in formal planning tasks, offering insights into their application and guiding future efforts in AI-driven planning paradigms.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20175
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Extensive Evaluation of PDDL Capabilities in off-the-shelf LLMs
Vyas, Kaustubh
Graux, Damien
Montella, Sébastien
Vougiouklis, Pavlos
Lai, Ruofei
Li, Keshuang
Ren, Yang
Pan, Jeff Z.
Artificial Intelligence
Computation and Language
In recent advancements, large language models (LLMs) have exhibited proficiency in code generation and chain-of-thought reasoning, laying the groundwork for tackling automatic formal planning tasks. This study evaluates the potential of LLMs to understand and generate Planning Domain Definition Language (PDDL), an essential representation in artificial intelligence planning. We conduct an extensive analysis across 20 distinct models spanning 7 major LLM families, both commercial and open-source. Our comprehensive evaluation sheds light on the zero-shot LLM capabilities of parsing, generating, and reasoning with PDDL. Our findings indicate that while some models demonstrate notable effectiveness in handling PDDL, others pose limitations in more complex scenarios requiring nuanced planning knowledge. These results highlight the promise and current limitations of LLMs in formal planning tasks, offering insights into their application and guiding future efforts in AI-driven planning paradigms.
title An Extensive Evaluation of PDDL Capabilities in off-the-shelf LLMs
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.20175