Prompt reinforcing for long-term planning of large language models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Hsien-Chin, Ruppik, Benjamin Matthias, van Niekerk, Carel, Shen, Chia-Hao, Heck, Michael, Lubis, Nurul, Vukovic, Renato, Feng, Shutong, Gašić, Milica
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911691296473088
author Lin, Hsien-Chin
Ruppik, Benjamin Matthias
van Niekerk, Carel
Shen, Chia-Hao
Heck, Michael
Lubis, Nurul
Vukovic, Renato
Feng, Shutong
Gašić, Milica
author_facet Lin, Hsien-Chin
Ruppik, Benjamin Matthias
van Niekerk, Carel
Shen, Chia-Hao
Heck, Michael
Lubis, Nurul
Vukovic, Renato
Feng, Shutong
Gašić, Milica
contents Large language models (LLMs) have achieved remarkable success in a wide range of natural language processing tasks and can be adapted through prompting. However, they remain suboptimal in multi-turn interactions, often relying on incorrect early assumptions and failing to track user goals over time, which makes such tasks particularly challenging. Prior works in dialogue systems have shown that long-term planning is essential for handling interactive tasks. In this work, we propose a prompt optimisation framework inspired by reinforcement learning, which enables such planning to take place by only modifying the task instruction prompt of the LLM-based agent. By generating turn-by-turn feedback and leveraging experience replay for prompt rewriting, our proposed method shows significant improvement in multi-turn tasks such as text-to-SQL and task-oriented dialogue. Moreover, it generalises across different LLM-based agents and can leverage diverse LLMs as meta-prompting agents. This warrants future research in reinforcement learning-inspired parameter-free optimisation methods.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05921
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prompt reinforcing for long-term planning of large language models
Lin, Hsien-Chin
Ruppik, Benjamin Matthias
van Niekerk, Carel
Shen, Chia-Hao
Heck, Michael
Lubis, Nurul
Vukovic, Renato
Feng, Shutong
Gašić, Milica
Computation and Language
Machine Learning
Large language models (LLMs) have achieved remarkable success in a wide range of natural language processing tasks and can be adapted through prompting. However, they remain suboptimal in multi-turn interactions, often relying on incorrect early assumptions and failing to track user goals over time, which makes such tasks particularly challenging. Prior works in dialogue systems have shown that long-term planning is essential for handling interactive tasks. In this work, we propose a prompt optimisation framework inspired by reinforcement learning, which enables such planning to take place by only modifying the task instruction prompt of the LLM-based agent. By generating turn-by-turn feedback and leveraging experience replay for prompt rewriting, our proposed method shows significant improvement in multi-turn tasks such as text-to-SQL and task-oriented dialogue. Moreover, it generalises across different LLM-based agents and can leverage diverse LLMs as meta-prompting agents. This warrants future research in reinforcement learning-inspired parameter-free optimisation methods.
title Prompt reinforcing for long-term planning of large language models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.05921