A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sarmah, Bhaskarjit, Dutta, Kriti, Grigoryan, Anna, Tiwari, Sachin, Pasquali, Stefano, Mehta, Dhagash
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916535118856192
author Sarmah, Bhaskarjit
Dutta, Kriti
Grigoryan, Anna
Tiwari, Sachin
Pasquali, Stefano
Mehta, Dhagash
author_facet Sarmah, Bhaskarjit
Dutta, Kriti
Grigoryan, Anna
Tiwari, Sachin
Pasquali, Stefano
Mehta, Dhagash
contents We argue that the Declarative Self-improving Python (DSPy) optimizers are a way to align the large language model (LLM) prompts and their evaluations to the human annotations. We present a comparative analysis of five teleprompter algorithms, namely, Cooperative Prompt Optimization (COPRO), Multi-Stage Instruction Prompt Optimization (MIPRO), BootstrapFewShot, BootstrapFewShot with Optuna, and K-Nearest Neighbor Few Shot, within the DSPy framework with respect to their ability to align with human evaluations. As a concrete example, we focus on optimizing the prompt to align hallucination detection (using LLM as a judge) to human annotated ground truth labels for a publicly available benchmark dataset. Our experiments demonstrate that optimized prompts can outperform various benchmark methods to detect hallucination, and certain telemprompters outperform the others in at least these experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15298
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation
Sarmah, Bhaskarjit
Dutta, Kriti
Grigoryan, Anna
Tiwari, Sachin
Pasquali, Stefano
Mehta, Dhagash
Computation and Language
Artificial Intelligence
Machine Learning
Statistical Finance
Methodology
We argue that the Declarative Self-improving Python (DSPy) optimizers are a way to align the large language model (LLM) prompts and their evaluations to the human annotations. We present a comparative analysis of five teleprompter algorithms, namely, Cooperative Prompt Optimization (COPRO), Multi-Stage Instruction Prompt Optimization (MIPRO), BootstrapFewShot, BootstrapFewShot with Optuna, and K-Nearest Neighbor Few Shot, within the DSPy framework with respect to their ability to align with human evaluations. As a concrete example, we focus on optimizing the prompt to align hallucination detection (using LLM as a judge) to human annotated ground truth labels for a publicly available benchmark dataset. Our experiments demonstrate that optimized prompts can outperform various benchmark methods to detect hallucination, and certain telemprompters outperform the others in at least these experiments.
title A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation
topic Computation and Language
Artificial Intelligence
Machine Learning
Statistical Finance
Methodology
url https://arxiv.org/abs/2412.15298