LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yuanchen, Verma, Saurabh, Lee, Justin, Xiong, Fangzhou, Zhang, Poppy, Awadelkarim, Amel, Chen, Xu, Yuan, Yubai, Hill, Shawndra
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917393507287040
author Wu, Yuanchen
Verma, Saurabh
Lee, Justin
Xiong, Fangzhou
Zhang, Poppy
Awadelkarim, Amel
Chen, Xu
Yuan, Yubai
Hill, Shawndra
author_facet Wu, Yuanchen
Verma, Saurabh
Lee, Justin
Xiong, Fangzhou
Zhang, Poppy
Awadelkarim, Amel
Chen, Xu
Yuan, Yubai
Hill, Shawndra
contents Large language models (LLMs) are highly sensitive to prompts, but most automatic prompt optimization (APO) methods assume access to ground-truth references (e.g., labeled validation data) that are costly to obtain. We propose the Prompt Duel Optimizer (PDO), a sample-efficient framework for label-free prompt optimization based on pairwise preference feedback from an LLM judge. PDO casts prompt selection as a dueling-bandit problem and combines (i) Double Thompson Sampling to prioritize informative comparisons under a fixed judge budget, with (ii) top-performer guided mutation to expand the candidate pool while pruning weak prompts. Experiments on BIG-bench Hard (BBH) and MS MARCO show that PDO consistently identifies stronger prompts than label-free baselines, while offering favorable quality--cost trade-offs under constrained comparison budgets.
format Preprint
id arxiv_https___arxiv_org_abs_2510_13907
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization
Wu, Yuanchen
Verma, Saurabh
Lee, Justin
Xiong, Fangzhou
Zhang, Poppy
Awadelkarim, Amel
Chen, Xu
Yuan, Yubai
Hill, Shawndra
Computation and Language
Machine Learning
Large language models (LLMs) are highly sensitive to prompts, but most automatic prompt optimization (APO) methods assume access to ground-truth references (e.g., labeled validation data) that are costly to obtain. We propose the Prompt Duel Optimizer (PDO), a sample-efficient framework for label-free prompt optimization based on pairwise preference feedback from an LLM judge. PDO casts prompt selection as a dueling-bandit problem and combines (i) Double Thompson Sampling to prioritize informative comparisons under a fixed judge budget, with (ii) top-performer guided mutation to expand the candidate pool while pruning weak prompts. Experiments on BIG-bench Hard (BBH) and MS MARCO show that PDO consistently identifies stronger prompts than label-free baselines, while offering favorable quality--cost trade-offs under constrained comparison budgets.
title LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.13907