2Mamba2Furious: Linear in Complexity, Competitive in Accuracy

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mongaras, Gabriel, Larson, Eric C.
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909045197111296
author Mongaras, Gabriel
Larson, Eric C.
author_facet Mongaras, Gabriel
Larson, Eric C.
contents Linear attention transformers have become a strong alternative to softmax attention due to their efficiency. However, linear attention tends to be less expressive and results in reduced accuracy compared to softmax attention. To bridge the accuracy gap between softmax attention and linear attention, we manipulate Mamba-2, a very strong linear attention variant. We first simplify Mamba-2 down to its most fundamental and important components, evaluating which specific choices make it most accurate. From this simplified Mamba variant (Mamba-2S), we improve the A-mask and increase the order of the hidden state, resulting in a method, which we call 2Mamba, that is nearly as accurate as softmax attention, yet much more memory efficient for long context lengths. We also investigate elements to Mamba-2 that help surpass softmax attention accuracy. Code is provided for all our experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2602_17363
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle 2Mamba2Furious: Linear in Complexity, Competitive in Accuracy
Mongaras, Gabriel
Larson, Eric C.
Machine Learning
I.2; I.2.6
Linear attention transformers have become a strong alternative to softmax attention due to their efficiency. However, linear attention tends to be less expressive and results in reduced accuracy compared to softmax attention. To bridge the accuracy gap between softmax attention and linear attention, we manipulate Mamba-2, a very strong linear attention variant. We first simplify Mamba-2 down to its most fundamental and important components, evaluating which specific choices make it most accurate. From this simplified Mamba variant (Mamba-2S), we improve the A-mask and increase the order of the hidden state, resulting in a method, which we call 2Mamba, that is nearly as accurate as softmax attention, yet much more memory efficient for long context lengths. We also investigate elements to Mamba-2 that help surpass softmax attention accuracy. Code is provided for all our experiments.
title 2Mamba2Furious: Linear in Complexity, Competitive in Accuracy
topic Machine Learning
I.2; I.2.6
url https://arxiv.org/abs/2602.17363