Non-myopic Generation of Language Models for Reasoning and Planning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Chang, Zhao, Haiteng, Zhang, Junlei, He, Junxian, Kong, Lingpeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914992797777920
author Ma, Chang
Zhao, Haiteng
Zhang, Junlei
He, Junxian
Kong, Lingpeng
author_facet Ma, Chang
Zhao, Haiteng
Zhang, Junlei
He, Junxian
Kong, Lingpeng
contents Large Language Models have demonstrated remarkable abilities in reasoning and planning by breaking down complex problems into sequential steps. Despite their success in various domains like mathematical problem-solving and coding, LLMs face challenges in ensuring reliable and optimal planning due to their inherent myopic nature of autoregressive decoding. This paper revisits LLM reasoning from an optimal-control perspective, proposing a novel method, Predictive-Decoding, that leverages Model Predictive Control to enhance planning accuracy. By re-weighting LLM distributions based on foresight trajectories, Predictive-Decoding aims to mitigate early errors and promote non-myopic planning. Our experiments show significant improvements in a wide range of tasks for math, coding, and agents. Furthermore, Predictive-Decoding demonstrates computational efficiency, outperforming search baselines with reduced computational resources. This study provides insights into optimizing LLM planning capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17195
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Non-myopic Generation of Language Models for Reasoning and Planning
Ma, Chang
Zhao, Haiteng
Zhang, Junlei
He, Junxian
Kong, Lingpeng
Artificial Intelligence
Computation and Language
Large Language Models have demonstrated remarkable abilities in reasoning and planning by breaking down complex problems into sequential steps. Despite their success in various domains like mathematical problem-solving and coding, LLMs face challenges in ensuring reliable and optimal planning due to their inherent myopic nature of autoregressive decoding. This paper revisits LLM reasoning from an optimal-control perspective, proposing a novel method, Predictive-Decoding, that leverages Model Predictive Control to enhance planning accuracy. By re-weighting LLM distributions based on foresight trajectories, Predictive-Decoding aims to mitigate early errors and promote non-myopic planning. Our experiments show significant improvements in a wide range of tasks for math, coding, and agents. Furthermore, Predictive-Decoding demonstrates computational efficiency, outperforming search baselines with reduced computational resources. This study provides insights into optimizing LLM planning capabilities.
title Non-myopic Generation of Language Models for Reasoning and Planning
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.17195