Enhancing LLM Planning Capabilities through Intrinsic Self-Critique

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bohnet, Bernd, Kamienny, Pierre-Alexandre, Sedghi, Hanie, Gorur, Dilan, Awasthi, Pranjal, Parisi, Aaron, Swersky, Kevin, Liu, Rosanne, Nova, Azade, Fiedel, Noah
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911346703990784
author Bohnet, Bernd
Kamienny, Pierre-Alexandre
Sedghi, Hanie
Gorur, Dilan
Awasthi, Pranjal
Parisi, Aaron
Swersky, Kevin
Liu, Rosanne
Nova, Azade
Fiedel, Noah
author_facet Bohnet, Bernd
Kamienny, Pierre-Alexandre
Sedghi, Hanie
Gorur, Dilan
Awasthi, Pranjal
Parisi, Aaron
Swersky, Kevin
Liu, Rosanne
Nova, Azade
Fiedel, Noah
contents We demonstrate an approach for LLMs to critique their \emph{own} answers with the goal of enhancing their performance that leads to significant improvements over established planning benchmarks. Despite the findings of earlier research that has cast doubt on the effectiveness of LLMs leveraging self critique methods, we show significant performance gains on planning datasets in the Blocksworld domain through intrinsic self-critique, without external source such as a verifier. We also demonstrate similar improvements on Logistics and Mini-grid datasets, exceeding strong baseline accuracies. We employ a few-shot learning technique and progressively extend it to a many-shot approach as our base method and demonstrate that it is possible to gain substantial improvement on top of this already competitive approach by employing an iterative process for correction and refinement. We illustrate how self-critique can significantly boost planning performance. Our empirical results present new state-of-the-art on the class of models considered, namely LLM model checkpoints from October 2024. Our primary focus lies on the method itself, demonstrating intrinsic self-improvement capabilities that are applicable regardless of the specific model version, and we believe that applying our method to more complex search techniques and more capable models will lead to even better performance.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24103
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing LLM Planning Capabilities through Intrinsic Self-Critique
Bohnet, Bernd
Kamienny, Pierre-Alexandre
Sedghi, Hanie
Gorur, Dilan
Awasthi, Pranjal
Parisi, Aaron
Swersky, Kevin
Liu, Rosanne
Nova, Azade
Fiedel, Noah
Machine Learning
Artificial Intelligence
We demonstrate an approach for LLMs to critique their \emph{own} answers with the goal of enhancing their performance that leads to significant improvements over established planning benchmarks. Despite the findings of earlier research that has cast doubt on the effectiveness of LLMs leveraging self critique methods, we show significant performance gains on planning datasets in the Blocksworld domain through intrinsic self-critique, without external source such as a verifier. We also demonstrate similar improvements on Logistics and Mini-grid datasets, exceeding strong baseline accuracies. We employ a few-shot learning technique and progressively extend it to a many-shot approach as our base method and demonstrate that it is possible to gain substantial improvement on top of this already competitive approach by employing an iterative process for correction and refinement. We illustrate how self-critique can significantly boost planning performance. Our empirical results present new state-of-the-art on the class of models considered, namely LLM model checkpoints from October 2024. Our primary focus lies on the method itself, demonstrating intrinsic self-improvement capabilities that are applicable regardless of the specific model version, and we believe that applying our method to more complex search techniques and more capable models will lead to even better performance.
title Enhancing LLM Planning Capabilities through Intrinsic Self-Critique
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2512.24103