Transformers, parallel computation, and logarithmic depth
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914679124656128 |
|---|---|
| author | Sanford, Clayton Hsu, Daniel Telgarsky, Matus |
| author_facet | Sanford, Clayton Hsu, Daniel Telgarsky, Matus |
| contents | We show that a constant number of self-attention layers can efficiently simulate, and be simulated by, a constant number of communication rounds of Massively Parallel Computation. As a consequence, we show that logarithmic depth is sufficient for transformers to solve basic computational tasks that cannot be efficiently solved by several other neural sequence models and sub-quadratic transformer approximations. We thus establish parallelism as a key distinguishing property of transformers. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_09268 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Transformers, parallel computation, and logarithmic depth Sanford, Clayton Hsu, Daniel Telgarsky, Matus Machine Learning We show that a constant number of self-attention layers can efficiently simulate, and be simulated by, a constant number of communication rounds of Massively Parallel Computation. As a consequence, we show that logarithmic depth is sufficient for transformers to solve basic computational tasks that cannot be efficiently solved by several other neural sequence models and sub-quadratic transformer approximations. We thus establish parallelism as a key distinguishing property of transformers. |
| title | Transformers, parallel computation, and logarithmic depth |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2402.09268 |