TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Nica, Andreea, Zakazov, Ivan, Baldwin, Nicolas Mario, Geng, Saibo, West, Robert
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909703977566208
author Nica, Andreea
Zakazov, Ivan
Baldwin, Nicolas Mario
Geng, Saibo
West, Robert
author_facet Nica, Andreea
Zakazov, Ivan
Baldwin, Nicolas Mario
Geng, Saibo
West, Robert
contents Prompt optimization improves the reasoning abilities of large language models (LLMs) without requiring parameter updates to the target model. Following heuristic-based "Think step by step" approaches, the field has evolved in two main directions: while one group of methods uses textual feedback to elicit improved prompts from general-purpose LLMs in a training-free way, a concurrent line of research relies on numerical rewards to train a special prompt model, tailored for providing optimal prompts to the target model. In this paper, we introduce the Textual Reward Prompt framework (TRPrompt), which unifies these approaches by directly incorporating textual feedback into training of the prompt model. Our framework does not require prior dataset collection and is being iteratively improved with the feedback on the generated prompts. When coupled with the capacity of an LLM to internalize the notion of what a "good" prompt is, the high-resolution signal provided by the textual rewards allows us to train a prompt model yielding state-of-the-art query-specific prompts for the problems from the challenging math datasets GSMHard and MATH.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18618
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards
Nica, Andreea
Zakazov, Ivan
Baldwin, Nicolas Mario
Geng, Saibo
West, Robert
Computation and Language
Machine Learning
Prompt optimization improves the reasoning abilities of large language models (LLMs) without requiring parameter updates to the target model. Following heuristic-based "Think step by step" approaches, the field has evolved in two main directions: while one group of methods uses textual feedback to elicit improved prompts from general-purpose LLMs in a training-free way, a concurrent line of research relies on numerical rewards to train a special prompt model, tailored for providing optimal prompts to the target model. In this paper, we introduce the Textual Reward Prompt framework (TRPrompt), which unifies these approaches by directly incorporating textual feedback into training of the prompt model. Our framework does not require prior dataset collection and is being iteratively improved with the feedback on the generated prompts. When coupled with the capacity of an LLM to internalize the notion of what a "good" prompt is, the high-resolution signal provided by the textual rewards allows us to train a prompt model yielding state-of-the-art query-specific prompts for the problems from the challenging math datasets GSMHard and MATH.
title TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2507.18618