MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map
Fuente:
arXiv
Saved in:
| Main Authors: | Chou, Yuhong, Yao, Man, Wang, Kexin, Pan, Yuqi, Zhu, Ruijie, Zhong, Yiran, Qiao, Yu, Wu, Jibin, Xu, Bo, Li, Guoqi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Linear Attention with Sparse State Expansion
by: Pan, Yuqi, et al.
Published: (2025)
by: Pan, Yuqi, et al.
Published: (2025)
ZeCO: Zero Communication Overhead Sequence Parallelism for Linear Attention
by: Chou, Yuhong, et al.
Published: (2025)
by: Chou, Yuhong, et al.
Published: (2025)
Integer-Valued Training and Spike-Driven Inference Spiking Neural Network for High-performance and Energy-efficient Object Detection
by: Luo, Xinhao, et al.
Published: (2024)
by: Luo, Xinhao, et al.
Published: (2024)
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
RSC-SNN: Exploring the Trade-off Between Adversarial Robustness and Accuracy in Spiking Neural Networks via Randomized Smoothing Coding
by: Wu, Keming, et al.
Published: (2024)
by: Wu, Keming, et al.
Published: (2024)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025)
by: Zuo, Yifei, et al.
Published: (2025)
Efficient 3D Recognition with Event-driven Spike Sparse Convolution
by: Qiu, Xuerui, et al.
Published: (2024)
by: Qiu, Xuerui, et al.
Published: (2024)
Agent Attention: On the Integration of Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2023)
by: Han, Dongchen, et al.
Published: (2023)
Scaling Spike-driven Transformer with Efficient Spike Firing Approximation Training
by: Yao, Man, et al.
Published: (2024)
by: Yao, Man, et al.
Published: (2024)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
by: Nishikawa, Naoki, et al.
Published: (2025)
by: Nishikawa, Naoki, et al.
Published: (2025)
Bridging the Divide: Reconsidering Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2024)
by: Han, Dongchen, et al.
Published: (2024)
High-Performance Temporal Reversible Spiking Neural Networks with $O(L)$ Training Memory and $O(1)$ Inference Cost
by: Hu, JiaKui, et al.
Published: (2024)
by: Hu, JiaKui, et al.
Published: (2024)
A Systematic Analysis of Hybrid Linear Attention
by: Wang, Dustin, et al.
Published: (2025)
by: Wang, Dustin, et al.
Published: (2025)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
Gated Attention Coding for Training High-performance and Efficient Spiking Neural Networks
by: Qiu, Xuerui, et al.
Published: (2023)
by: Qiu, Xuerui, et al.
Published: (2023)
Linear Attention Sequence Parallelism
by: Sun, Weigao, et al.
Published: (2024)
by: Sun, Weigao, et al.
Published: (2024)
Scalable Autoregressive Image Generation with Mamba
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
Softmax Linear Attention: Reclaiming Global Competition
by: Xu, Mingwei, et al.
Published: (2026)
by: Xu, Mingwei, et al.
Published: (2026)
IML-Spikeformer: Input-aware Multi-Level Spiking Transformer for Speech Processing
by: Song, Zeyang, et al.
Published: (2025)
by: Song, Zeyang, et al.
Published: (2025)
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
by: Lin, Max Qiushi, et al.
Published: (2025)
by: Lin, Max Qiushi, et al.
Published: (2025)
Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
by: Mitchell, Rupert, et al.
Published: (2025)
by: Mitchell, Rupert, et al.
Published: (2025)
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
by: Zhang, Michael, et al.
Published: (2024)
by: Zhang, Michael, et al.
Published: (2024)
Elucidating the Design Space of Decay in Linear Attention
by: Qin, Zhen, et al.
Published: (2025)
by: Qin, Zhen, et al.
Published: (2025)
Parallel Training in Spiking Neural Networks
by: Huang, Yanbin, et al.
Published: (2026)
by: Huang, Yanbin, et al.
Published: (2026)
SpikingBrain: Spiking Brain-inspired Large Models
by: Pan, Yuqi, et al.
Published: (2025)
by: Pan, Yuqi, et al.
Published: (2025)
SpikeVoice: High-Quality Text-to-Speech Via Efficient Spiking Neural Network
by: Wang, Kexin, et al.
Published: (2024)
by: Wang, Kexin, et al.
Published: (2024)
Softmax-free Linear Transformers
by: Lu, Jiachen, et al.
Published: (2022)
by: Lu, Jiachen, et al.
Published: (2022)
In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention
by: He, Jianliang, et al.
Published: (2025)
by: He, Jianliang, et al.
Published: (2025)
On the Invariants of Softmax Attention
by: Lee, Wonsuk
Published: (2026)
by: Lee, Wonsuk
Published: (2026)
Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning
by: Xie, Zixuan, et al.
Published: (2026)
by: Xie, Zixuan, et al.
Published: (2026)
Targeted Protein Degradation by Repurposing Transmembrane E3 Ubiquitin Ligases
by: Jibin Cui, et al.
Published: (2026)
by: Jibin Cui, et al.
Published: (2026)
Scalable-Softmax Is Superior for Attention
by: Nakanishi, Ken M.
Published: (2025)
by: Nakanishi, Ken M.
Published: (2025)
Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
by: Boursier, Etienne, et al.
Published: (2025)
by: Boursier, Etienne, et al.
Published: (2025)
Optimal Approximation of Elliptic Problems by Linear and Nonlinear Mappings II
by: Dahlke, Stephan, et al.
Published: (2006)
by: Dahlke, Stephan, et al.
Published: (2006)
Optimal Approximation of Elliptic Problems by Linear and Nonlinear Mappings I
by: Dahlke, Stephan, et al.
Published: (2006)
by: Dahlke, Stephan, et al.
Published: (2006)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025)
Spike-driven Transformer V2: Meta Spiking Neural Network Architecture Inspiring the Design of Next-generation Neuromorphic Chips
by: Yao, Man, et al.
Published: (2024)
by: Yao, Man, et al.
Published: (2024)
SoLA-Vision: Fine-grained Layer-wise Linear Softmax Hybrid Attention
by: Li, Ruibang, et al.
Published: (2026)
by: Li, Ruibang, et al.
Published: (2026)
Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
NDCG-Consistent Softmax Approximation with Accelerated Convergence
by: Pu, Yuanhao, et al.
Published: (2025)
by: Pu, Yuanhao, et al.
Published: (2025)
Similar Items
-
Scaling Linear Attention with Sparse State Expansion
by: Pan, Yuqi, et al.
Published: (2025) -
ZeCO: Zero Communication Overhead Sequence Parallelism for Linear Attention
by: Chou, Yuhong, et al.
Published: (2025) -
Integer-Valued Training and Spike-Driven Inference Spiking Neural Network for High-performance and Energy-efficient Object Detection
by: Luo, Xinhao, et al.
Published: (2024) -
Universal Approximation with Softmax Attention
by: Hu, Jerry Yao-Chieh, et al.
Published: (2025) -
RSC-SNN: Exploring the Trade-off Between Adversarial Robustness and Accuracy in Spiking Neural Networks via Randomized Smoothing Coding
by: Wu, Keming, et al.
Published: (2024)