Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bigelow, Eric, Wurgaft, Daniel, Wang, YingQiao, Goodman, Noah, Ullman, Tomer, Tanaka, Hidenori, Lubana, Ekdeep Singh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914387298615296
author Bigelow, Eric
Wurgaft, Daniel
Wang, YingQiao
Goodman, Noah
Ullman, Tomer
Tanaka, Hidenori
Lubana, Ekdeep Singh
author_facet Bigelow, Eric
Wurgaft, Daniel
Wang, YingQiao
Goodman, Noah
Ullman, Tomer
Tanaka, Hidenori
Lubana, Ekdeep Singh
contents Large language models (LLMs) can be controlled at inference time through prompts (in-context learning) and internal activations (activation steering). Different accounts have been proposed to explain these methods, yet their common goal of controlling model behavior raises the question of whether these seemingly disparate methodologies can be seen as specific instances of a broader framework. Motivated by this, we develop a unifying, predictive account of LLM control from a Bayesian perspective. Specifically, we posit that both context- and activation-based interventions impact model behavior by altering its belief in latent concepts: steering operates by changing concept priors, while in-context learning leads to an accumulation of evidence. This results in a closed-form Bayesian model that is highly predictive of LLM behavior across context- and activation-based interventions in a set of domains inspired by prior work on many-shot in-context learning. This model helps us explain prior empirical phenomena - e.g., sigmoidal learning curves as in-context evidence accumulates - while predicting novel ones - e.g., additivity of both interventions in log-belief space, which results in distinct phases such that sudden and dramatic behavioral shifts can be induced by slightly changing intervention controls. Taken together, this work offers a unified account of prompt-based and activation-based control of LLM behavior, and a methodology for empirically predicting the effects of these interventions.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00617
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
Bigelow, Eric
Wurgaft, Daniel
Wang, YingQiao
Goodman, Noah
Ullman, Tomer
Tanaka, Hidenori
Lubana, Ekdeep Singh
Machine Learning
Artificial Intelligence
Computation and Language
Large language models (LLMs) can be controlled at inference time through prompts (in-context learning) and internal activations (activation steering). Different accounts have been proposed to explain these methods, yet their common goal of controlling model behavior raises the question of whether these seemingly disparate methodologies can be seen as specific instances of a broader framework. Motivated by this, we develop a unifying, predictive account of LLM control from a Bayesian perspective. Specifically, we posit that both context- and activation-based interventions impact model behavior by altering its belief in latent concepts: steering operates by changing concept priors, while in-context learning leads to an accumulation of evidence. This results in a closed-form Bayesian model that is highly predictive of LLM behavior across context- and activation-based interventions in a set of domains inspired by prior work on many-shot in-context learning. This model helps us explain prior empirical phenomena - e.g., sigmoidal learning curves as in-context evidence accumulates - while predicting novel ones - e.g., additivity of both interventions in log-belief space, which results in distinct phases such that sudden and dramatic behavioral shifts can be induced by slightly changing intervention controls. Taken together, this work offers a unified account of prompt-based and activation-based control of LLM behavior, and a methodology for empirically predicting the effects of these interventions.
title Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2511.00617