Saved in:
Bibliographic Details
Main Authors: Sanford, Clayton, Hsu, Daniel, Telgarsky, Matus
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2402.09268
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914679124656128
author Sanford, Clayton
Hsu, Daniel
Telgarsky, Matus
author_facet Sanford, Clayton
Hsu, Daniel
Telgarsky, Matus
contents We show that a constant number of self-attention layers can efficiently simulate, and be simulated by, a constant number of communication rounds of Massively Parallel Computation. As a consequence, we show that logarithmic depth is sufficient for transformers to solve basic computational tasks that cannot be efficiently solved by several other neural sequence models and sub-quadratic transformer approximations. We thus establish parallelism as a key distinguishing property of transformers.
format Preprint
id arxiv_https___arxiv_org_abs_2402_09268
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Transformers, parallel computation, and logarithmic depth
Sanford, Clayton
Hsu, Daniel
Telgarsky, Matus
Machine Learning
We show that a constant number of self-attention layers can efficiently simulate, and be simulated by, a constant number of communication rounds of Massively Parallel Computation. As a consequence, we show that logarithmic depth is sufficient for transformers to solve basic computational tasks that cannot be efficiently solved by several other neural sequence models and sub-quadratic transformer approximations. We thus establish parallelism as a key distinguishing property of transformers.
title Transformers, parallel computation, and logarithmic depth
topic Machine Learning
url https://arxiv.org/abs/2402.09268