PLHF: Prompt Optimization with Few-Shot Human Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Chun-Pai, Zheng, Kan, Lin, Shou-De
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912372770209792
author Yang, Chun-Pai
Zheng, Kan
Lin, Shou-De
author_facet Yang, Chun-Pai
Zheng, Kan
Lin, Shou-De
contents Automatic prompt optimization frameworks are developed to obtain suitable prompts for large language models (LLMs) with respect to desired output quality metrics. Although existing approaches can handle conventional tasks such as fixed-solution question answering, defining the metric becomes complicated when the output quality cannot be easily assessed by comparisons with standard golden samples. Consequently, optimizing the prompts effectively and efficiently without a clear metric becomes a critical challenge. To address the issue, we present PLHF (which stands for "P"rompt "L"earning with "H"uman "F"eedback), a few-shot prompt optimization framework inspired by the well-known RLHF technique. Different from naive strategies, PLHF employs a specific evaluator module acting as the metric to estimate the output quality. PLHF requires only a single round of human feedback to complete the entire prompt optimization process. Empirical results on both public and industrial datasets show that PLHF outperforms prior output grading strategies for LLM prompt optimizations.
format Preprint
id arxiv_https___arxiv_org_abs_2505_07886
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PLHF: Prompt Optimization with Few-Shot Human Feedback
Yang, Chun-Pai
Zheng, Kan
Lin, Shou-De
Computation and Language
Artificial Intelligence
Automatic prompt optimization frameworks are developed to obtain suitable prompts for large language models (LLMs) with respect to desired output quality metrics. Although existing approaches can handle conventional tasks such as fixed-solution question answering, defining the metric becomes complicated when the output quality cannot be easily assessed by comparisons with standard golden samples. Consequently, optimizing the prompts effectively and efficiently without a clear metric becomes a critical challenge. To address the issue, we present PLHF (which stands for "P"rompt "L"earning with "H"uman "F"eedback), a few-shot prompt optimization framework inspired by the well-known RLHF technique. Different from naive strategies, PLHF employs a specific evaluator module acting as the metric to estimate the output quality. PLHF requires only a single round of human feedback to complete the entire prompt optimization process. Empirical results on both public and industrial datasets show that PLHF outperforms prior output grading strategies for LLM prompt optimizations.
title PLHF: Prompt Optimization with Few-Shot Human Feedback
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.07886