ACING: Actor-Critic for Instruction Learning in Black-Box LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kharrat, Salma, Fourati, Fares, Canini, Marco
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908518696615936
author Kharrat, Salma
Fourati, Fares
Canini, Marco
author_facet Kharrat, Salma
Fourati, Fares
Canini, Marco
contents The effectiveness of Large Language Models (LLMs) in solving tasks depends significantly on the quality of their instructions, which often require substantial human effort to craft. This underscores the need for automated instruction optimization. However, optimizing instructions is particularly challenging when working with black-box LLMs, where model parameters and gradients are inaccessible. We introduce ACING, an actor-critic reinforcement learning framework that formulates instruction optimization as a stateless, continuous-action problem, enabling exploration of infinite instruction spaces using only black-box feedback. ACING automatically discovers prompts that outperform human-written prompts in 76% of instruction-induction tasks, with gains of up to 33 points and a 10-point median improvement over the best automatic baseline in 33 tasks spanning instruction-induction, summarization, and chain-of-thought reasoning. Extensive ablations highlight its robustness and efficiency. An implementation of ACING is available at https://github.com/salmakh1/ACING.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12736
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
Kharrat, Salma
Fourati, Fares
Canini, Marco
Computation and Language
Artificial Intelligence
Machine Learning
Systems and Control
Optimization and Control
The effectiveness of Large Language Models (LLMs) in solving tasks depends significantly on the quality of their instructions, which often require substantial human effort to craft. This underscores the need for automated instruction optimization. However, optimizing instructions is particularly challenging when working with black-box LLMs, where model parameters and gradients are inaccessible. We introduce ACING, an actor-critic reinforcement learning framework that formulates instruction optimization as a stateless, continuous-action problem, enabling exploration of infinite instruction spaces using only black-box feedback. ACING automatically discovers prompts that outperform human-written prompts in 76% of instruction-induction tasks, with gains of up to 33 points and a 10-point median improvement over the best automatic baseline in 33 tasks spanning instruction-induction, summarization, and chain-of-thought reasoning. Extensive ablations highlight its robustness and efficiency. An implementation of ACING is available at https://github.com/salmakh1/ACING.
title ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
topic Computation and Language
Artificial Intelligence
Machine Learning
Systems and Control
Optimization and Control
url https://arxiv.org/abs/2411.12736