On the Optimization and Generalization of Multi-head Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deora, Puneesh, Ghaderi, Rouzbeh, Taheri, Hossein, Thrampoulidis, Christos
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909346149957632
author Deora, Puneesh
Ghaderi, Rouzbeh
Taheri, Hossein
Thrampoulidis, Christos
author_facet Deora, Puneesh
Ghaderi, Rouzbeh
Taheri, Hossein
Thrampoulidis, Christos
contents The training and generalization dynamics of the Transformer's core mechanism, namely the Attention mechanism, remain under-explored. Besides, existing analyses primarily focus on single-head attention. Inspired by the demonstrated benefits of overparameterization when training fully-connected networks, we investigate the potential optimization and generalization advantages of using multiple attention heads. Towards this goal, we derive convergence and generalization guarantees for gradient-descent training of a single-layer multi-head self-attention model, under a suitable realizability condition on the data. We then establish primitive conditions on the initialization that ensure realizability holds. Finally, we demonstrate that these conditions are satisfied for a simple tokenized-mixture model. We expect the analysis can be extended to various data-model and architecture variations.
format Preprint
id arxiv_https___arxiv_org_abs_2310_12680
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle On the Optimization and Generalization of Multi-head Attention
Deora, Puneesh
Ghaderi, Rouzbeh
Taheri, Hossein
Thrampoulidis, Christos
Machine Learning
Optimization and Control
The training and generalization dynamics of the Transformer's core mechanism, namely the Attention mechanism, remain under-explored. Besides, existing analyses primarily focus on single-head attention. Inspired by the demonstrated benefits of overparameterization when training fully-connected networks, we investigate the potential optimization and generalization advantages of using multiple attention heads. Towards this goal, we derive convergence and generalization guarantees for gradient-descent training of a single-layer multi-head self-attention model, under a suitable realizability condition on the data. We then establish primitive conditions on the initialization that ensure realizability holds. Finally, we demonstrate that these conditions are satisfied for a simple tokenized-mixture model. We expect the analysis can be extended to various data-model and architecture variations.
title On the Optimization and Generalization of Multi-head Attention
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2310.12680