Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteur principal: Musat, Tiberiu
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911235539206144
author Musat, Tiberiu
author_facet Musat, Tiberiu
contents In this paper, I introduce the retrieval problem, a simple yet common reasoning task that can be solved only by transformers with a minimum number of layers, which grows logarithmically with the input size. I empirically show that large language models can solve the task under different prompting formulations without any fine-tuning. To understand how transformers solve the retrieval problem, I train several transformers on a minimal formulation. Successful learning occurs only under the presence of an implicit curriculum. I uncover the learned mechanisms by studying the attention maps in the trained transformers. I also study the training process, uncovering that attention heads always emerge in a specific sequence guided by the implicit curriculum.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12118
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
Musat, Tiberiu
Machine Learning
Computation and Language
In this paper, I introduce the retrieval problem, a simple yet common reasoning task that can be solved only by transformers with a minimum number of layers, which grows logarithmically with the input size. I empirically show that large language models can solve the task under different prompting formulations without any fine-tuning. To understand how transformers solve the retrieval problem, I train several transformers on a minimal formulation. Successful learning occurs only under the presence of an implicit curriculum. I uncover the learned mechanisms by studying the attention maps in the trained transformers. I also study the training process, uncovering that attention heads always emerge in a specific sequence guided by the implicit curriculum.
title Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2411.12118