On the Optimization and Generalization of Multi-head Attention
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909346149957632 |
|---|---|
| author | Deora, Puneesh Ghaderi, Rouzbeh Taheri, Hossein Thrampoulidis, Christos |
| author_facet | Deora, Puneesh Ghaderi, Rouzbeh Taheri, Hossein Thrampoulidis, Christos |
| contents | The training and generalization dynamics of the Transformer's core mechanism, namely the Attention mechanism, remain under-explored. Besides, existing analyses primarily focus on single-head attention. Inspired by the demonstrated benefits of overparameterization when training fully-connected networks, we investigate the potential optimization and generalization advantages of using multiple attention heads. Towards this goal, we derive convergence and generalization guarantees for gradient-descent training of a single-layer multi-head self-attention model, under a suitable realizability condition on the data. We then establish primitive conditions on the initialization that ensure realizability holds. Finally, we demonstrate that these conditions are satisfied for a simple tokenized-mixture model. We expect the analysis can be extended to various data-model and architecture variations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2310_12680 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | On the Optimization and Generalization of Multi-head Attention Deora, Puneesh Ghaderi, Rouzbeh Taheri, Hossein Thrampoulidis, Christos Machine Learning Optimization and Control The training and generalization dynamics of the Transformer's core mechanism, namely the Attention mechanism, remain under-explored. Besides, existing analyses primarily focus on single-head attention. Inspired by the demonstrated benefits of overparameterization when training fully-connected networks, we investigate the potential optimization and generalization advantages of using multiple attention heads. Towards this goal, we derive convergence and generalization guarantees for gradient-descent training of a single-layer multi-head self-attention model, under a suitable realizability condition on the data. We then establish primitive conditions on the initialization that ensure realizability holds. Finally, we demonstrate that these conditions are satisfied for a simple tokenized-mixture model. We expect the analysis can be extended to various data-model and architecture variations. |
| title | On the Optimization and Generalization of Multi-head Attention |
| topic | Machine Learning Optimization and Control |
| url | https://arxiv.org/abs/2310.12680 |