Amortizing intractable inference in large language models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Edward J., Jain, Moksh, Elmoznino, Eric, Kaddar, Younesse, Lajoie, Guillaume, Bengio, Yoshua, Malkin, Nikolay
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910366595809280
author Hu, Edward J.
Jain, Moksh
Elmoznino, Eric
Kaddar, Younesse
Lajoie, Guillaume
Bengio, Yoshua
Malkin, Nikolay
author_facet Hu, Edward J.
Jain, Moksh
Elmoznino, Eric
Kaddar, Younesse
Lajoie, Guillaume
Bengio, Yoshua
Malkin, Nikolay
contents Autoregressive large language models (LLMs) compress knowledge from their training data through next-token conditional distributions. This limits tractable querying of this knowledge to start-to-end autoregressive sampling. However, many tasks of interest -- including sequence continuation, infilling, and other forms of constrained generation -- involve sampling from intractable posterior distributions. We address this limitation by using amortized Bayesian inference to sample from these intractable posteriors. Such amortization is algorithmically achieved by fine-tuning LLMs via diversity-seeking reinforcement learning algorithms: generative flow networks (GFlowNets). We empirically demonstrate that this distribution-matching paradigm of LLM fine-tuning can serve as an effective alternative to maximum-likelihood training and reward-maximizing policy optimization. As an important application, we interpret chain-of-thought reasoning as a latent variable modeling problem and demonstrate that our approach enables data-efficient adaptation of LLMs to tasks that require multi-step rationalization and tool use.
format Preprint
id arxiv_https___arxiv_org_abs_2310_04363
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Amortizing intractable inference in large language models
Hu, Edward J.
Jain, Moksh
Elmoznino, Eric
Kaddar, Younesse
Lajoie, Guillaume
Bengio, Yoshua
Malkin, Nikolay
Machine Learning
Computation and Language
Autoregressive large language models (LLMs) compress knowledge from their training data through next-token conditional distributions. This limits tractable querying of this knowledge to start-to-end autoregressive sampling. However, many tasks of interest -- including sequence continuation, infilling, and other forms of constrained generation -- involve sampling from intractable posterior distributions. We address this limitation by using amortized Bayesian inference to sample from these intractable posteriors. Such amortization is algorithmically achieved by fine-tuning LLMs via diversity-seeking reinforcement learning algorithms: generative flow networks (GFlowNets). We empirically demonstrate that this distribution-matching paradigm of LLM fine-tuning can serve as an effective alternative to maximum-likelihood training and reward-maximizing policy optimization. As an important application, we interpret chain-of-thought reasoning as a latent variable modeling problem and demonstrate that our approach enables data-efficient adaptation of LLMs to tasks that require multi-step rationalization and tool use.
title Amortizing intractable inference in large language models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2310.04363