Language models and brains align due to more than next-word prediction and word-level information

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Merlin, Gabriele, Toneva, Mariya
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913527955980288
author Merlin, Gabriele
Toneva, Mariya
author_facet Merlin, Gabriele
Toneva, Mariya
contents Pretrained language models have been shown to significantly predict brain recordings of people comprehending language. Recent work suggests that the prediction of the next word is a key mechanism that contributes to this alignment. What is not yet understood is whether prediction of the next word is necessary for this observed alignment or simply sufficient, and whether there are other shared mechanisms or information that are similarly important. In this work, we take a step towards understanding the reasons for brain alignment via two simple perturbations in popular pretrained language models. These perturbations help us design contrasts that can control for different types of information. By contrasting the brain alignment of these differently perturbed models, we show that improvements in alignment with brain recordings are due to more than improvements in next-word prediction and word-level information.
format Preprint
id arxiv_https___arxiv_org_abs_2212_00596
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Language models and brains align due to more than next-word prediction and word-level information
Merlin, Gabriele
Toneva, Mariya
Computation and Language
Neurons and Cognition
Pretrained language models have been shown to significantly predict brain recordings of people comprehending language. Recent work suggests that the prediction of the next word is a key mechanism that contributes to this alignment. What is not yet understood is whether prediction of the next word is necessary for this observed alignment or simply sufficient, and whether there are other shared mechanisms or information that are similarly important. In this work, we take a step towards understanding the reasons for brain alignment via two simple perturbations in popular pretrained language models. These perturbations help us design contrasts that can control for different types of information. By contrasting the brain alignment of these differently perturbed models, we show that improvements in alignment with brain recordings are due to more than improvements in next-word prediction and word-level information.
title Language models and brains align due to more than next-word prediction and word-level information
topic Computation and Language
Neurons and Cognition
url https://arxiv.org/abs/2212.00596