Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bortoletto, Matteo, Ruhdorfer, Constantin, Shi, Lei, Bulling, Andreas
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913846109667328
author Bortoletto, Matteo
Ruhdorfer, Constantin
Shi, Lei
Bulling, Andreas
author_facet Bortoletto, Matteo
Ruhdorfer, Constantin
Shi, Lei
Bulling, Andreas
contents Despite growing interest in Theory of Mind (ToM) tasks for evaluating language models (LMs), little is known about how LMs internally represent mental states of self and others. Understanding these internal mechanisms is critical - not only to move beyond surface-level performance, but also for model alignment and safety, where subtle misattributions of mental states may go undetected in generated outputs. In this work, we present the first systematic investigation of belief representations in LMs by probing models across different scales, training regimens, and prompts - using control tasks to rule out confounds. Our experiments provide evidence that both model size and fine-tuning substantially improve LMs' internal representations of others' beliefs, which are structured - not mere by-products of spurious correlations - yet brittle to prompt variations. Crucially, we show that these representations can be strengthened: targeted edits to model activations can correct wrong ToM inferences.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17513
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models
Bortoletto, Matteo
Ruhdorfer, Constantin
Shi, Lei
Bulling, Andreas
Computation and Language
Artificial Intelligence
Despite growing interest in Theory of Mind (ToM) tasks for evaluating language models (LMs), little is known about how LMs internally represent mental states of self and others. Understanding these internal mechanisms is critical - not only to move beyond surface-level performance, but also for model alignment and safety, where subtle misattributions of mental states may go undetected in generated outputs. In this work, we present the first systematic investigation of belief representations in LMs by probing models across different scales, training regimens, and prompts - using control tasks to rule out confounds. Our experiments provide evidence that both model size and fine-tuning substantially improve LMs' internal representations of others' beliefs, which are structured - not mere by-products of spurious correlations - yet brittle to prompt variations. Crucially, we show that these representations can be strengthened: targeted edits to model activations can correct wrong ToM inferences.
title Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2406.17513