A completely uniform transformer for parity
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909448558084096 |
|---|---|
| author | Kozachinskiy, Alexander Steifer, Tomasz |
| author_facet | Kozachinskiy, Alexander Steifer, Tomasz |
| contents | We construct a 3-layer constant-dimension transformer, recognizing the parity language, where neither parameter matrices nor the positional encoding depend on the input length. This improves upon a construction of Chiang and Cholak who use a positional encoding, depending on the input length (but their construction has 2 layers). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_02535 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A completely uniform transformer for parity Kozachinskiy, Alexander Steifer, Tomasz Machine Learning Artificial Intelligence We construct a 3-layer constant-dimension transformer, recognizing the parity language, where neither parameter matrices nor the positional encoding depend on the input length. This improves upon a construction of Chiang and Cholak who use a positional encoding, depending on the input length (but their construction has 2 layers). |
| title | A completely uniform transformer for parity |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2501.02535 |