The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sarfati, Raphaël, Bigelow, Eric, Wurgaft, Daniel, Boppana, Siddharth, Merullo, Jack, Geiger, Atticus, Lewis, Owen, McGrath, Tom, Lubana, Ekdeep Singh
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909015608393728
author Sarfati, Raphaël
Bigelow, Eric
Wurgaft, Daniel
Boppana, Siddharth
Merullo, Jack
Geiger, Atticus
Lewis, Owen
McGrath, Tom
Lubana, Ekdeep Singh
author_facet Sarfati, Raphaël
Bigelow, Eric
Wurgaft, Daniel
Boppana, Siddharth
Merullo, Jack
Geiger, Atticus
Lewis, Owen
McGrath, Tom
Lubana, Ekdeep Singh
contents Large language models (LLMs) form implicit beliefs (posteriors over latent variables) from prompts, but we lack a mechanistic account of how these beliefs are encoded in representation space, how they update with new evidence, and how interventions reshape them. We study a controlled setting in which Llama-3.2 infers the parameters of a normal distribution from in-context samples. We show that parameter posteriors are encoded as curved manifolds in representation space, and trace how they evolve along the prompt. Standard linear steering moves representations off-manifold, inducing unintended, coupled changes, whereas geometry-aware methods preserve the target belief family. Our work demonstrates an example of linear field probing (LFP) as a principled approach to tile the data manifold and make interventions that respect the underlying geometry. Our results suggest that LLM beliefs are inherently geometric objects, and that globally linear representations are often inadequate abstractions.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02315
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
Sarfati, Raphaël
Bigelow, Eric
Wurgaft, Daniel
Boppana, Siddharth
Merullo, Jack
Geiger, Atticus
Lewis, Owen
McGrath, Tom
Lubana, Ekdeep Singh
Computation and Language
Large language models (LLMs) form implicit beliefs (posteriors over latent variables) from prompts, but we lack a mechanistic account of how these beliefs are encoded in representation space, how they update with new evidence, and how interventions reshape them. We study a controlled setting in which Llama-3.2 infers the parameters of a normal distribution from in-context samples. We show that parameter posteriors are encoded as curved manifolds in representation space, and trace how they evolve along the prompt. Standard linear steering moves representations off-manifold, inducing unintended, coupled changes, whereas geometry-aware methods preserve the target belief family. Our work demonstrates an example of linear field probing (LFP) as a principled approach to tile the data manifold and make interventions that respect the underlying geometry. Our results suggest that LLM beliefs are inherently geometric objects, and that globally linear representations are often inadequate abstractions.
title The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
topic Computation and Language
url https://arxiv.org/abs/2602.02315