Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2402.09268 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914679124656128 |
|---|---|
| author | Sanford, Clayton Hsu, Daniel Telgarsky, Matus |
| author_facet | Sanford, Clayton Hsu, Daniel Telgarsky, Matus |
| contents | We show that a constant number of self-attention layers can efficiently simulate, and be simulated by, a constant number of communication rounds of Massively Parallel Computation. As a consequence, we show that logarithmic depth is sufficient for transformers to solve basic computational tasks that cannot be efficiently solved by several other neural sequence models and sub-quadratic transformer approximations. We thus establish parallelism as a key distinguishing property of transformers. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_09268 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Transformers, parallel computation, and logarithmic depth Sanford, Clayton Hsu, Daniel Telgarsky, Matus Machine Learning We show that a constant number of self-attention layers can efficiently simulate, and be simulated by, a constant number of communication rounds of Massively Parallel Computation. As a consequence, we show that logarithmic depth is sufficient for transformers to solve basic computational tasks that cannot be efficiently solved by several other neural sequence models and sub-quadratic transformer approximations. We thus establish parallelism as a key distinguishing property of transformers. |
| title | Transformers, parallel computation, and logarithmic depth |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2402.09268 |