Understanding Factual Recall in Transformers via Associative Memories

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nichani, Eshaan, Lee, Jason D., Bietti, Alberto
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910734957412352
author Nichani, Eshaan
Lee, Jason D.
Bietti, Alberto
author_facet Nichani, Eshaan
Lee, Jason D.
Bietti, Alberto
contents Large language models have demonstrated an impressive ability to perform factual recall. Prior work has found that transformers trained on factual recall tasks can store information at a rate proportional to their parameter count. In our work, we show that shallow transformers can use a combination of associative memories to obtain such near optimal storage capacity. We begin by proving that the storage capacities of both linear and MLP associative memories scale linearly with parameter count. We next introduce a synthetic factual recall task, and prove that a transformer with a single layer of self-attention followed by an MLP can obtain 100% accuracy on the task whenever either the total number of self-attention parameters or MLP parameters scales (up to log factors) linearly with the number of facts. In particular, the transformer can trade off between using the value matrices or the MLP as an associative memory to store the dataset of facts. We complement these expressivity results with an analysis of the gradient flow trajectory of a simplified linear attention model trained on our factual recall task, where we show that the model exhibits sequential learning behavior.
format Preprint
id arxiv_https___arxiv_org_abs_2412_06538
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Understanding Factual Recall in Transformers via Associative Memories
Nichani, Eshaan
Lee, Jason D.
Bietti, Alberto
Machine Learning
Computation and Language
Information Theory
Large language models have demonstrated an impressive ability to perform factual recall. Prior work has found that transformers trained on factual recall tasks can store information at a rate proportional to their parameter count. In our work, we show that shallow transformers can use a combination of associative memories to obtain such near optimal storage capacity. We begin by proving that the storage capacities of both linear and MLP associative memories scale linearly with parameter count. We next introduce a synthetic factual recall task, and prove that a transformer with a single layer of self-attention followed by an MLP can obtain 100% accuracy on the task whenever either the total number of self-attention parameters or MLP parameters scales (up to log factors) linearly with the number of facts. In particular, the transformer can trade off between using the value matrices or the MLP as an associative memory to store the dataset of facts. We complement these expressivity results with an analysis of the gradient flow trajectory of a simplified linear attention model trained on our factual recall task, where we show that the model exhibits sequential learning behavior.
title Understanding Factual Recall in Transformers via Associative Memories
topic Machine Learning
Computation and Language
Information Theory
url https://arxiv.org/abs/2412.06538