Salvato in:
Dettagli Bibliografici
Autori principali: Kumar, Saurabh, Buckman, Jacob, Gelada, Carles, Zhang, Sean
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:https://arxiv.org/abs/2503.03269
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908348713009152
author Kumar, Saurabh
Buckman, Jacob
Gelada, Carles
Zhang, Sean
author_facet Kumar, Saurabh
Buckman, Jacob
Gelada, Carles
Zhang, Sean
contents Transformers with linear attention offer significant computational advantages over softmax-based transformers but often suffer from degraded performance. The symmetric power (sympow) transformer, a particular type of linear transformer, addresses some of this performance gap by leveraging symmetric tensor embeddings, achieving comparable performance to softmax transformers. However, the finite capacity of the recurrent state in sympow transformers limits their ability to retain information, leading to performance degradation when scaling the training or evaluation context length. To address this issue, we propose the conformal-sympow transformer, which dynamically frees up capacity using data-dependent multiplicative gating and adaptively stores information using data-dependent rotary embeddings. Preliminary experiments on the LongCrawl64 dataset demonstrate that conformal-sympow overcomes the limitations of sympow transformers, achieving robust performance across scaled training and evaluation contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03269
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Conformal Transformations for Symmetric Power Transformers
Kumar, Saurabh
Buckman, Jacob
Gelada, Carles
Zhang, Sean
Machine Learning
Artificial Intelligence
Transformers with linear attention offer significant computational advantages over softmax-based transformers but often suffer from degraded performance. The symmetric power (sympow) transformer, a particular type of linear transformer, addresses some of this performance gap by leveraging symmetric tensor embeddings, achieving comparable performance to softmax transformers. However, the finite capacity of the recurrent state in sympow transformers limits their ability to retain information, leading to performance degradation when scaling the training or evaluation context length. To address this issue, we propose the conformal-sympow transformer, which dynamically frees up capacity using data-dependent multiplicative gating and adaptively stores information using data-dependent rotary embeddings. Preliminary experiments on the LongCrawl64 dataset demonstrate that conformal-sympow overcomes the limitations of sympow transformers, achieving robust performance across scaled training and evaluation contexts.
title Conformal Transformations for Symmetric Power Transformers
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2503.03269