Bayesian scaling laws for in-context learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arora, Aryaman, Jurafsky, Dan, Potts, Christopher, Goodman, Noah D.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912597628944384
author Arora, Aryaman
Jurafsky, Dan
Potts, Christopher
Goodman, Noah D.
author_facet Arora, Aryaman
Jurafsky, Dan
Potts, Christopher
Goodman, Noah D.
contents In-context learning (ICL) is a powerful technique for getting language models to perform complex tasks with no training updates. Prior work has established strong correlations between the number of in-context examples provided and the accuracy of the model's predictions. In this paper, we seek to explain this correlation by showing that ICL approximates a Bayesian learner. This perspective gives rise to a novel Bayesian scaling law for ICL. In experiments with \mbox{GPT-2} models of different sizes, our scaling law matches existing scaling laws in accuracy while also offering interpretable terms for task priors, learning efficiency, and per-example probabilities. To illustrate the analytic power that such interpretable scaling laws provide, we report on controlled synthetic dataset experiments designed to inform real-world studies of safety alignment. In our experimental protocol, we use SFT or DPO to suppress an unwanted existing model capability and then use ICL to try to bring that capability back (many-shot jailbreaking). We then study ICL on real-world instruction-tuned LLMs using capabilities benchmarks as well as a new many-shot jailbreaking dataset. In all cases, Bayesian scaling laws accurately predict the conditions under which ICL will cause suppressed behaviors to reemerge, which sheds light on the ineffectiveness of post-training at increasing LLM safety.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16531
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Bayesian scaling laws for in-context learning
Arora, Aryaman
Jurafsky, Dan
Potts, Christopher
Goodman, Noah D.
Computation and Language
Artificial Intelligence
Formal Languages and Automata Theory
Machine Learning
I.2.7
In-context learning (ICL) is a powerful technique for getting language models to perform complex tasks with no training updates. Prior work has established strong correlations between the number of in-context examples provided and the accuracy of the model's predictions. In this paper, we seek to explain this correlation by showing that ICL approximates a Bayesian learner. This perspective gives rise to a novel Bayesian scaling law for ICL. In experiments with \mbox{GPT-2} models of different sizes, our scaling law matches existing scaling laws in accuracy while also offering interpretable terms for task priors, learning efficiency, and per-example probabilities. To illustrate the analytic power that such interpretable scaling laws provide, we report on controlled synthetic dataset experiments designed to inform real-world studies of safety alignment. In our experimental protocol, we use SFT or DPO to suppress an unwanted existing model capability and then use ICL to try to bring that capability back (many-shot jailbreaking). We then study ICL on real-world instruction-tuned LLMs using capabilities benchmarks as well as a new many-shot jailbreaking dataset. In all cases, Bayesian scaling laws accurately predict the conditions under which ICL will cause suppressed behaviors to reemerge, which sheds light on the ineffectiveness of post-training at increasing LLM safety.
title Bayesian scaling laws for in-context learning
topic Computation and Language
Artificial Intelligence
Formal Languages and Automata Theory
Machine Learning
I.2.7
url https://arxiv.org/abs/2410.16531