Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Lei, Bruna, Joan, Bietti, Alberto
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:https://arxiv.org/abs/2406.03068
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915184585474048
author Chen, Lei
Bruna, Joan
Bietti, Alberto
author_facet Chen, Lei
Bruna, Joan
Bietti, Alberto
contents Large language models have been successful at tasks involving basic forms of in-context reasoning, such as generating coherent language, as well as storing vast amounts of knowledge. At the core of the Transformer architecture behind such models are feed-forward and attention layers, which are often associated to knowledge and reasoning, respectively. In this paper, we study this distinction empirically and theoretically in a controlled synthetic setting where certain next-token predictions involve both distributional and in-context information. We find that feed-forward layers tend to learn simple distributional associations such as bigrams, while attention layers focus on in-context reasoning. Our theoretical analysis identifies the noise in the gradients as a key factor behind this discrepancy. Finally, we illustrate how similar disparities emerge in pre-trained models through ablations on the Pythia model family on simple reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_03068
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers
Chen, Lei
Bruna, Joan
Bietti, Alberto
Machine Learning
Artificial Intelligence
Computation and Language
Large language models have been successful at tasks involving basic forms of in-context reasoning, such as generating coherent language, as well as storing vast amounts of knowledge. At the core of the Transformer architecture behind such models are feed-forward and attention layers, which are often associated to knowledge and reasoning, respectively. In this paper, we study this distinction empirically and theoretically in a controlled synthetic setting where certain next-token predictions involve both distributional and in-context information. We find that feed-forward layers tend to learn simple distributional associations such as bigrams, while attention layers focus on in-context reasoning. Our theoretical analysis identifies the noise in the gradients as a key factor behind this discrepancy. Finally, we illustrate how similar disparities emerge in pre-trained models through ablations on the Pythia model family on simple reasoning tasks.
title Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2406.03068