Memorization Capacity of Multi-Head Attention in Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mahdavi, Sadegh, Liao, Renjie, Thrampoulidis, Christos
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910349882556416
author Mahdavi, Sadegh
Liao, Renjie
Thrampoulidis, Christos
author_facet Mahdavi, Sadegh
Liao, Renjie
Thrampoulidis, Christos
contents Transformers have become the go-to architecture for language and vision tasks, yet their theoretical properties, especially memorization capacity, remain elusive. This paper investigates the memorization abilities of multi-head attention mechanisms, examining how many example sequences they can memorize, as a function of the number of heads and sequence length. Motivated by experimental findings on vision transformers, we introduce novel assumptions about the linear independence of input data, distinct from the commonly used general-position assumption. Under these assumptions, we demonstrate that an attention layer with $H$ heads, dimension $d$, and context size $n < d$, featuring $Θ(Hd^2)$ parameters, can memorize $Ω(Hn)$ examples. Our analysis sheds light on how different attention heads handle various example sequences, aided by the softmax operator's saturation property. We validate our findings through experiments on synthetic data.
format Preprint
id arxiv_https___arxiv_org_abs_2306_02010
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Memorization Capacity of Multi-Head Attention in Transformers
Mahdavi, Sadegh
Liao, Renjie
Thrampoulidis, Christos
Machine Learning
Transformers have become the go-to architecture for language and vision tasks, yet their theoretical properties, especially memorization capacity, remain elusive. This paper investigates the memorization abilities of multi-head attention mechanisms, examining how many example sequences they can memorize, as a function of the number of heads and sequence length. Motivated by experimental findings on vision transformers, we introduce novel assumptions about the linear independence of input data, distinct from the commonly used general-position assumption. Under these assumptions, we demonstrate that an attention layer with $H$ heads, dimension $d$, and context size $n < d$, featuring $Θ(Hd^2)$ parameters, can memorize $Ω(Hn)$ examples. Our analysis sheds light on how different attention heads handle various example sequences, aided by the softmax operator's saturation property. We validate our findings through experiments on synthetic data.
title Memorization Capacity of Multi-Head Attention in Transformers
topic Machine Learning
url https://arxiv.org/abs/2306.02010