Linear Log-Normal Attention with Unbiased Concentration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nahshan, Yury, Kampeas, Joseph, Haleva, Emir
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914691294429184
author Nahshan, Yury
Kampeas, Joseph
Haleva, Emir
author_facet Nahshan, Yury
Kampeas, Joseph
Haleva, Emir
contents Transformer models have achieved remarkable results in a wide range of applications. However, their scalability is hampered by the quadratic time and memory complexity of the self-attention mechanism concerning the sequence length. This limitation poses a substantial obstacle when dealing with long documents or high-resolution images. In this work, we study the self-attention mechanism by analyzing the distribution of the attention matrix and its concentration ability. Furthermore, we propose instruments to measure these quantities and introduce a novel self-attention mechanism, Linear Log-Normal Attention, designed to emulate the distribution and concentration behavior of the original self-attention. Our experimental results on popular natural language benchmarks reveal that our proposed Linear Log-Normal Attention outperforms other linearized attention alternatives, offering a promising avenue for enhancing the scalability of transformer models.
format Preprint
id arxiv_https___arxiv_org_abs_2311_13541
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Linear Log-Normal Attention with Unbiased Concentration
Nahshan, Yury
Kampeas, Joseph
Haleva, Emir
Machine Learning
Artificial Intelligence
I.7.0; G.3
Transformer models have achieved remarkable results in a wide range of applications. However, their scalability is hampered by the quadratic time and memory complexity of the self-attention mechanism concerning the sequence length. This limitation poses a substantial obstacle when dealing with long documents or high-resolution images. In this work, we study the self-attention mechanism by analyzing the distribution of the attention matrix and its concentration ability. Furthermore, we propose instruments to measure these quantities and introduce a novel self-attention mechanism, Linear Log-Normal Attention, designed to emulate the distribution and concentration behavior of the original self-attention. Our experimental results on popular natural language benchmarks reveal that our proposed Linear Log-Normal Attention outperforms other linearized attention alternatives, offering a promising avenue for enhancing the scalability of transformer models.
title Linear Log-Normal Attention with Unbiased Concentration
topic Machine Learning
Artificial Intelligence
I.7.0; G.3
url https://arxiv.org/abs/2311.13541