PMPO: Probabilistic Metric Prompt Optimization for Small and Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Chenzhuo, Liu, Ziqian, Wang, Xinda, Lu, Junting, Ruan, Chaoyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914045325475840
author Zhao, Chenzhuo
Liu, Ziqian
Wang, Xinda
Lu, Junting
Ruan, Chaoyi
author_facet Zhao, Chenzhuo
Liu, Ziqian
Wang, Xinda
Lu, Junting
Ruan, Chaoyi
contents Prompt optimization is a practical and widely applicable alternative to fine tuning for improving large language model performance. Yet many existing methods evaluate candidate prompts by sampling full outputs, often coupled with self critique or human annotated preferences, which limits scalability, especially for smaller models or models that are not instruction tuned. We present PMPO (Probabilistic Metric Prompt Optimization), a unified framework that uses token level cross entropy as a direct, lightweight evaluation signal. PMPO locates low quality prompt segments via a masking based analysis and iteratively rewrites them to propose improved variants. Crucially, during evaluation, PMPO selects among variants by minimizing loss in a single forward pass, eliminating output sampling and human or judge based scoring for selection while still using standard generation only to propose rewrites. This unified, loss based strategy supports both supervised and preference based tasks. Across model sizes and datasets, PMPO outperforms prior prompt optimizers: it achieves the highest average accuracy on BBH, performs strongly on GSM8K and AQUA RAT, and raises AlpacaEval 2.0 win rates by over 19 points. These results demonstrate PMPO's effectiveness, efficiency, and broad applicability.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16307
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PMPO: Probabilistic Metric Prompt Optimization for Small and Large Language Models
Zhao, Chenzhuo
Liu, Ziqian
Wang, Xinda
Lu, Junting
Ruan, Chaoyi
Computation and Language
Artificial Intelligence
Machine Learning
Prompt optimization is a practical and widely applicable alternative to fine tuning for improving large language model performance. Yet many existing methods evaluate candidate prompts by sampling full outputs, often coupled with self critique or human annotated preferences, which limits scalability, especially for smaller models or models that are not instruction tuned. We present PMPO (Probabilistic Metric Prompt Optimization), a unified framework that uses token level cross entropy as a direct, lightweight evaluation signal. PMPO locates low quality prompt segments via a masking based analysis and iteratively rewrites them to propose improved variants. Crucially, during evaluation, PMPO selects among variants by minimizing loss in a single forward pass, eliminating output sampling and human or judge based scoring for selection while still using standard generation only to propose rewrites. This unified, loss based strategy supports both supervised and preference based tasks. Across model sizes and datasets, PMPO outperforms prior prompt optimizers: it achieves the highest average accuracy on BBH, performs strongly on GSM8K and AQUA RAT, and raises AlpacaEval 2.0 win rates by over 19 points. These results demonstrate PMPO's effectiveness, efficiency, and broad applicability.
title PMPO: Probabilistic Metric Prompt Optimization for Small and Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.16307