Learning How Hard to Think: Input-Adaptive Allocation of LM Computation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Damani, Mehul, Shenfeld, Idan, Peng, Andi, Bobu, Andreea, Andreas, Jacob
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929530065649664
author Damani, Mehul
Shenfeld, Idan
Peng, Andi
Bobu, Andreea
Andreas, Jacob
author_facet Damani, Mehul
Shenfeld, Idan
Peng, Andi
Bobu, Andreea
Andreas, Jacob
contents Computationally intensive decoding procedures--including search, reranking, and self-critique--can improve the quality of language model (LM) outputs in problems spanning code generation, numerical reasoning, and dialog. Existing work typically applies the same decoding procedure for every input to an LM. But not all inputs require the same amount of computation to process. Can we allocate decoding computation adaptively, using more resources to answer questions whose answers will be harder to compute? We present an approach that predicts the distribution of rewards given an input and computation budget, then allocates additional computation to inputs for which it is predicted to be most useful. We apply this approach in two decoding procedures: first, an adaptive best-of-k procedure that dynamically selects the number of samples to generate as input to a reranker; second, a routing procedure that dynamically responds to a query using a decoding procedure that is expensive but accurate, or one that is cheaper but less capable. Across a suite of programming, mathematics, and dialog tasks, we show that accurate computation-allocation procedures can be learned, and reduce computation by up to 50% at no cost to response quality, or improve quality by up to 10% at a fixed computational budget.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04707
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
Damani, Mehul
Shenfeld, Idan
Peng, Andi
Bobu, Andreea
Andreas, Jacob
Machine Learning
Artificial Intelligence
Computation and Language
Computationally intensive decoding procedures--including search, reranking, and self-critique--can improve the quality of language model (LM) outputs in problems spanning code generation, numerical reasoning, and dialog. Existing work typically applies the same decoding procedure for every input to an LM. But not all inputs require the same amount of computation to process. Can we allocate decoding computation adaptively, using more resources to answer questions whose answers will be harder to compute? We present an approach that predicts the distribution of rewards given an input and computation budget, then allocates additional computation to inputs for which it is predicted to be most useful. We apply this approach in two decoding procedures: first, an adaptive best-of-k procedure that dynamically selects the number of samples to generate as input to a reranker; second, a routing procedure that dynamically responds to a query using a decoding procedure that is expensive but accurate, or one that is cheaper but less capable. Across a suite of programming, mathematics, and dialog tasks, we show that accurate computation-allocation procedures can be learned, and reduce computation by up to 50% at no cost to response quality, or improve quality by up to 10% at a fixed computational budget.
title Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.04707