MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chou, Yuhong, Yao, Man, Wang, Kexin, Pan, Yuqi, Zhu, Ruijie, Zhong, Yiran, Qiao, Yu, Wu, Jibin, Xu, Bo, Li, Guoqi
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916483492216832
author Chou, Yuhong
Yao, Man
Wang, Kexin
Pan, Yuqi
Zhu, Ruijie
Zhong, Yiran
Qiao, Yu
Wu, Jibin
Xu, Bo
Li, Guoqi
author_facet Chou, Yuhong
Yao, Man
Wang, Kexin
Pan, Yuqi
Zhu, Ruijie
Zhong, Yiran
Qiao, Yu
Wu, Jibin
Xu, Bo
Li, Guoqi
contents Various linear complexity models, such as Linear Transformer (LinFormer), State Space Model (SSM), and Linear RNN (LinRNN), have been proposed to replace the conventional softmax attention in Transformer structures. However, the optimal design of these linear models is still an open question. In this work, we attempt to answer this question by finding the best linear approximation to softmax attention from a theoretical perspective. We start by unifying existing linear complexity models as the linear attention form and then identify three conditions for the optimal linear attention design: 1) Dynamic memory ability; 2) Static approximation ability; 3) Least parameter approximation. We find that none of the current linear models meet all three conditions, resulting in suboptimal performance. Instead, we propose Meta Linear Attention (MetaLA) as a solution that satisfies these conditions. Our experiments on Multi-Query Associative Recall (MQAR) task, language modeling, image classification, and Long-Range Arena (LRA) benchmark demonstrate that MetaLA is more effective than the existing linear models.
format Preprint
id arxiv_https___arxiv_org_abs_2411_10741
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map
Chou, Yuhong
Yao, Man
Wang, Kexin
Pan, Yuqi
Zhu, Ruijie
Zhong, Yiran
Qiao, Yu
Wu, Jibin
Xu, Bo
Li, Guoqi
Machine Learning
Artificial Intelligence
Various linear complexity models, such as Linear Transformer (LinFormer), State Space Model (SSM), and Linear RNN (LinRNN), have been proposed to replace the conventional softmax attention in Transformer structures. However, the optimal design of these linear models is still an open question. In this work, we attempt to answer this question by finding the best linear approximation to softmax attention from a theoretical perspective. We start by unifying existing linear complexity models as the linear attention form and then identify three conditions for the optimal linear attention design: 1) Dynamic memory ability; 2) Static approximation ability; 3) Least parameter approximation. We find that none of the current linear models meet all three conditions, resulting in suboptimal performance. Instead, we propose Meta Linear Attention (MetaLA) as a solution that satisfies these conditions. Our experiments on Multi-Query Associative Recall (MQAR) task, language modeling, image classification, and Long-Range Arena (LRA) benchmark demonstrate that MetaLA is more effective than the existing linear models.
title MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2411.10741