Adaptive Memory Decay for Log-Linear Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Amin, Yaxita, Li, Helen Zichen, Zhang, Mengfan, Ayhan, Samet |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
E2Former-V2: On-the-Fly Equivariant Attention with Linear Activation Memory
by: Huang, Lin, et al.
Published: (2026)
by: Huang, Lin, et al.
Published: (2026)
Linear Log-Normal Attention with Unbiased Concentration
by: Nahshan, Yury, et al.
Published: (2023)
by: Nahshan, Yury, et al.
Published: (2023)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025)
by: Zuo, Yifei, et al.
Published: (2025)
Stochastic Sparse Attention for Memory-Bound Inference
by: Lee, Kyle, et al.
Published: (2026)
by: Lee, Kyle, et al.
Published: (2026)
Optimal Linear Decay Learning Rate Schedules and Further Refinements
by: Defazio, Aaron, et al.
Published: (2023)
by: Defazio, Aaron, et al.
Published: (2023)
FedProphet: Memory-Efficient Federated Adversarial Training via Robust and Consistent Cascade Learning
by: Tang, Minxue, et al.
Published: (2024)
by: Tang, Minxue, et al.
Published: (2024)
Regret Bounds and Reinforcement Learning Exploration of EXP-based Algorithms
by: Xu, Mengfan, et al.
Published: (2020)
by: Xu, Mengfan, et al.
Published: (2020)
Exact Linear Attention
by: Ou, Weinuo
Published: (2026)
by: Ou, Weinuo
Published: (2026)
Kaczmarz Linear Attention
by: Zou, Jiaxuan, et al.
Published: (2026)
by: Zou, Jiaxuan, et al.
Published: (2026)
Optimal Decay Spectra for Linear Recurrences
by: Cao, Yang
Published: (2026)
by: Cao, Yang
Published: (2026)
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
by: Li, Yingcong, et al.
Published: (2024)
by: Li, Yingcong, et al.
Published: (2024)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
by: Hu, Wenjie, et al.
Published: (2025)
by: Hu, Wenjie, et al.
Published: (2025)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
by: Ildiz, M. Emrullah, et al.
Published: (2024)
by: Ildiz, M. Emrullah, et al.
Published: (2024)
SmartMixed: A Two-Phase Training Strategy for Adaptive Activation Function Learning in Neural Networks
by: Omidvar, Amin
Published: (2025)
by: Omidvar, Amin
Published: (2025)
State Rank Dynamics in Linear Attention LLMs
by: Sun, Ao, et al.
Published: (2026)
by: Sun, Ao, et al.
Published: (2026)
Enhancing Linear Attention with Residual Learning
by: Lai, Xunhao, et al.
Published: (2025)
by: Lai, Xunhao, et al.
Published: (2025)
Hyperbolic Hypergraph Neural Networks for Multi-Relational Knowledge Hypergraph Representation
by: Li, Mengfan, et al.
Published: (2024)
by: Li, Mengfan, et al.
Published: (2024)
Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
by: Xu, Huangyu, et al.
Published: (2026)
by: Xu, Huangyu, et al.
Published: (2026)
AFD-STA: Adaptive Filtering Denoising with Spatiotemporal Attention for Chaotic System Prediction
by: Gong, Chunlin, et al.
Published: (2025)
by: Gong, Chunlin, et al.
Published: (2025)
Localized Observation Abstraction Using Piecewise Linear Spatial Decay for Reinforcement Learning in Combat Simulations
by: Black, Scotty, et al.
Published: (2024)
by: Black, Scotty, et al.
Published: (2024)
Higher-order Linear Attention
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Adaptive Locally Linear Embedding
by: Goli, Ali, et al.
Published: (2025)
by: Goli, Ali, et al.
Published: (2025)
KVBuffer: IO-aware Serving for Linear Attention
by: Zou, Longwei, et al.
Published: (2026)
by: Zou, Longwei, et al.
Published: (2026)
Sequential Policy Gradient for Adaptive Hyperparameter Optimization
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
Linear Attention is Enough in Spatial-Temporal Forecasting
by: Ning, Xinyu
Published: (2024)
by: Ning, Xinyu
Published: (2024)
Linear Attention for Efficient Bidirectional Sequence Modeling
by: Afzal, Arshia, et al.
Published: (2025)
by: Afzal, Arshia, et al.
Published: (2025)
Choosing How to Remember: Adaptive Memory Structures for LLM Agents
by: Lu, Mingfei, et al.
Published: (2026)
by: Lu, Mingfei, et al.
Published: (2026)
Echo State Transformer: Attention Over Finite Memories
by: Bendi-Ouis, Yannis, et al.
Published: (2025)
by: Bendi-Ouis, Yannis, et al.
Published: (2025)
Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
by: Beck, Maximilian, et al.
Published: (2025)
by: Beck, Maximilian, et al.
Published: (2025)
Beyond Linearity in Attention Projections: The Case for Nonlinear Queries
by: Karbevski, Marko
Published: (2026)
by: Karbevski, Marko
Published: (2026)
ZeroS: Zero-Sum Linear Attention for Efficient Transformers
by: Lu, Jiecheng, et al.
Published: (2026)
by: Lu, Jiecheng, et al.
Published: (2026)
Efficient Linear Attention for Multivariate Time Series Modeling via Entropy Equality
by: Zhang, Mingtao, et al.
Published: (2025)
by: Zhang, Mingtao, et al.
Published: (2025)
Superiority of Multi-Head Attention in In-Context Linear Regression
by: Cui, Yingqian, et al.
Published: (2024)
by: Cui, Yingqian, et al.
Published: (2024)
FAME: Adaptive Functional Attention with Expert Routing for Function-on-Function Regression
by: Gao, Yifei, et al.
Published: (2025)
by: Gao, Yifei, et al.
Published: (2025)
Mechanics of Next Token Prediction with Self-Attention
by: Li, Yingcong, et al.
Published: (2024)
by: Li, Yingcong, et al.
Published: (2024)
Provably Efficient Algorithm for Best Scoring Rule Identification in Online Principal-Agent Information Acquisition
by: Wang, Zichen, et al.
Published: (2025)
by: Wang, Zichen, et al.
Published: (2025)
MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map
by: Chou, Yuhong, et al.
Published: (2024)
by: Chou, Yuhong, et al.
Published: (2024)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
by: Joshi, Sahil, et al.
Published: (2025)
by: Joshi, Sahil, et al.
Published: (2025)
High-Fidelity Industrial Crash Dynamics Prediction via Geometry-Aware Operator Learning with Memory-Efficient Low-Rank Attention
by: Akhare, Deepak, et al.
Published: (2026)
by: Akhare, Deepak, et al.
Published: (2026)
Similar Items
-
E2Former-V2: On-the-Fly Equivariant Attention with Linear Activation Memory
by: Huang, Lin, et al.
Published: (2026) -
Linear Log-Normal Attention with Unbiased Concentration
by: Nahshan, Yury, et al.
Published: (2023) -
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
by: Li, Yingcong, et al.
Published: (2025) -
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025) -
Stochastic Sparse Attention for Memory-Bound Inference
by: Lee, Kyle, et al.
Published: (2026)