LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917393507287040 |
|---|---|
| author | Wu, Yuanchen Verma, Saurabh Lee, Justin Xiong, Fangzhou Zhang, Poppy Awadelkarim, Amel Chen, Xu Yuan, Yubai Hill, Shawndra |
| author_facet | Wu, Yuanchen Verma, Saurabh Lee, Justin Xiong, Fangzhou Zhang, Poppy Awadelkarim, Amel Chen, Xu Yuan, Yubai Hill, Shawndra |
| contents | Large language models (LLMs) are highly sensitive to prompts, but most automatic prompt optimization (APO) methods assume access to ground-truth references (e.g., labeled validation data) that are costly to obtain. We propose the Prompt Duel Optimizer (PDO), a sample-efficient framework for label-free prompt optimization based on pairwise preference feedback from an LLM judge. PDO casts prompt selection as a dueling-bandit problem and combines (i) Double Thompson Sampling to prioritize informative comparisons under a fixed judge budget, with (ii) top-performer guided mutation to expand the candidate pool while pruning weak prompts. Experiments on BIG-bench Hard (BBH) and MS MARCO show that PDO consistently identifies stronger prompts than label-free baselines, while offering favorable quality--cost trade-offs under constrained comparison budgets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_13907 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization Wu, Yuanchen Verma, Saurabh Lee, Justin Xiong, Fangzhou Zhang, Poppy Awadelkarim, Amel Chen, Xu Yuan, Yubai Hill, Shawndra Computation and Language Machine Learning Large language models (LLMs) are highly sensitive to prompts, but most automatic prompt optimization (APO) methods assume access to ground-truth references (e.g., labeled validation data) that are costly to obtain. We propose the Prompt Duel Optimizer (PDO), a sample-efficient framework for label-free prompt optimization based on pairwise preference feedback from an LLM judge. PDO casts prompt selection as a dueling-bandit problem and combines (i) Double Thompson Sampling to prioritize informative comparisons under a fixed judge budget, with (ii) top-performer guided mutation to expand the candidate pool while pruning weak prompts. Experiments on BIG-bench Hard (BBH) and MS MARCO show that PDO consistently identifies stronger prompts than label-free baselines, while offering favorable quality--cost trade-offs under constrained comparison budgets. |
| title | LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2510.13907 |