Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Opsahl-Ong, Krista, Ryan, Michael J, Purtell, Josh, Broman, David, Potts, Christopher, Zaharia, Matei, Khattab, Omar
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913535273992192
author Opsahl-Ong, Krista
Ryan, Michael J
Purtell, Josh
Broman, David
Potts, Christopher
Zaharia, Matei
Khattab, Omar
author_facet Opsahl-Ong, Krista
Ryan, Michael J
Purtell, Josh
Broman, David
Potts, Christopher
Zaharia, Matei
Khattab, Omar
contents Language Model Programs, i.e. sophisticated pipelines of modular language model (LM) calls, are increasingly advancing NLP tasks, but they require crafting prompts that are jointly effective for all modules. We study prompt optimization for LM programs, i.e. how to update these prompts to maximize a downstream metric without access to module-level labels or gradients. To make this tractable, we factorize our problem into optimizing the free-form instructions and few-shot demonstrations of every module and introduce several strategies to craft task-grounded instructions and navigate credit assignment across modules. Our strategies include (i) program- and data-aware techniques for proposing effective instructions, (ii) a stochastic mini-batch evaluation function for learning a surrogate model of our objective, and (iii) a meta-optimization procedure in which we refine how LMs construct proposals over time. Using these insights we develop MIPRO, a novel algorithm for optimizing LM programs. MIPRO outperforms baseline optimizers on five of seven diverse multi-stage LM programs using a best-in-class open-source model (Llama-3-8B), by as high as 13% accuracy. We have released our new optimizers and benchmark in DSPy at http://dspy.ai
format Preprint
id arxiv_https___arxiv_org_abs_2406_11695
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs
Opsahl-Ong, Krista
Ryan, Michael J
Purtell, Josh
Broman, David
Potts, Christopher
Zaharia, Matei
Khattab, Omar
Computation and Language
Artificial Intelligence
Machine Learning
Language Model Programs, i.e. sophisticated pipelines of modular language model (LM) calls, are increasingly advancing NLP tasks, but they require crafting prompts that are jointly effective for all modules. We study prompt optimization for LM programs, i.e. how to update these prompts to maximize a downstream metric without access to module-level labels or gradients. To make this tractable, we factorize our problem into optimizing the free-form instructions and few-shot demonstrations of every module and introduce several strategies to craft task-grounded instructions and navigate credit assignment across modules. Our strategies include (i) program- and data-aware techniques for proposing effective instructions, (ii) a stochastic mini-batch evaluation function for learning a surrogate model of our objective, and (iii) a meta-optimization procedure in which we refine how LMs construct proposals over time. Using these insights we develop MIPRO, a novel algorithm for optimizing LM programs. MIPRO outperforms baseline optimizers on five of seven diverse multi-stage LM programs using a best-in-class open-source model (Llama-3-8B), by as high as 13% accuracy. We have released our new optimizers and benchmark in DSPy at http://dspy.ai
title Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.11695