The Key to State Reduction in Linear Attention: A Rank-based Perspective
Fuente:
arXiv
Guardado en:
| Autores principales: | Nazari, Philipp, Rusch, T. Konstantin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Curious Case of In-Training Compression of State Space Models
por: Chahine, Makram, et al.
Publicado: (2025)
por: Chahine, Makram, et al.
Publicado: (2025)
Oscillatory State-Space Models
por: Rusch, T. Konstantin, et al.
Publicado: (2024)
por: Rusch, T. Konstantin, et al.
Publicado: (2024)
Learning to Dissipate Energy in Oscillatory State-Space Models
por: Boyer, Jared, et al.
Publicado: (2025)
por: Boyer, Jared, et al.
Publicado: (2025)
Low-Pass Flow Matching
por: Ruscio, Francesco M., et al.
Publicado: (2026)
por: Ruscio, Francesco M., et al.
Publicado: (2026)
State Rank Dynamics in Linear Attention LLMs
por: Sun, Ao, et al.
Publicado: (2026)
por: Sun, Ao, et al.
Publicado: (2026)
Low-Rank Key Value Attention
por: O'Neill, James, et al.
Publicado: (2026)
por: O'Neill, James, et al.
Publicado: (2026)
Quantifying Memory Use in Reinforcement Learning with Temporal Range
por: Lafuente-Mercado, Rodney, et al.
Publicado: (2025)
por: Lafuente-Mercado, Rodney, et al.
Publicado: (2025)
Low Stein Discrepancy via Message-Passing Monte Carlo
por: Kirk, Nathan, et al.
Publicado: (2025)
por: Kirk, Nathan, et al.
Publicado: (2025)
Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
por: Boursier, Etienne, et al.
Publicado: (2025)
por: Boursier, Etienne, et al.
Publicado: (2025)
Relaxed Equivariance via Multitask Learning
por: Elhag, Ahmed A., et al.
Publicado: (2024)
por: Elhag, Ahmed A., et al.
Publicado: (2024)
LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
por: Shahbazi, Ashkan, et al.
Publicado: (2025)
por: Shahbazi, Ashkan, et al.
Publicado: (2025)
LoLA: Low-Rank Linear Attention With Sparse Caching
por: McDermott, Luke, et al.
Publicado: (2025)
por: McDermott, Luke, et al.
Publicado: (2025)
Rank Reduction Autoencoders
por: Mounayer, Jad, et al.
Publicado: (2024)
por: Mounayer, Jad, et al.
Publicado: (2024)
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
por: Wu, Ziyang, et al.
Publicado: (2024)
por: Wu, Ziyang, et al.
Publicado: (2024)
Variational Rank Reduction Autoencoders
por: Mounayer, Jad, et al.
Publicado: (2025)
por: Mounayer, Jad, et al.
Publicado: (2025)
Power-based Partial Attention: Bridging Linear-Complexity and Full Attention
por: Huang, Yufeng
Publicado: (2026)
por: Huang, Yufeng
Publicado: (2026)
On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention
por: Ro, Yeonju, et al.
Publicado: (2025)
por: Ro, Yeonju, et al.
Publicado: (2025)
Message-Passing Monte Carlo: Generating low-discrepancy point sets via Graph Neural Networks
por: Rusch, T. Konstantin, et al.
Publicado: (2024)
por: Rusch, T. Konstantin, et al.
Publicado: (2024)
Scaling Linear Attention with Sparse State Expansion
por: Pan, Yuqi, et al.
Publicado: (2025)
por: Pan, Yuqi, et al.
Publicado: (2025)
On the Benefits of Rank in Attention Layers
por: Amsel, Noah, et al.
Publicado: (2024)
por: Amsel, Noah, et al.
Publicado: (2024)
Neural Low-Discrepancy Sequences
por: Van Huffel, Michael Etienne, et al.
Publicado: (2025)
por: Van Huffel, Michael Etienne, et al.
Publicado: (2025)
Multi-Head Low-Rank Attention
por: Liu, Songtao, et al.
Publicado: (2026)
por: Liu, Songtao, et al.
Publicado: (2026)
Log-Linear Attention
por: Guo, Han, et al.
Publicado: (2025)
por: Guo, Han, et al.
Publicado: (2025)
Kalman Linear Attention: Parallel Bayesian Filtering For Efficient Language Modelling and State Tracking
por: Shaj, Vaisakh, et al.
Publicado: (2026)
por: Shaj, Vaisakh, et al.
Publicado: (2026)
How does over-squashing affect the power of GNNs?
por: Di Giovanni, Francesco, et al.
Publicado: (2023)
por: Di Giovanni, Francesco, et al.
Publicado: (2023)
Low-Rank Tensor Decompositions for the Theory of Neural Networks
por: Borsoi, Ricardo, et al.
Publicado: (2025)
por: Borsoi, Ricardo, et al.
Publicado: (2025)
Causal Attention with Lookahead Keys
por: Song, Zhuoqing, et al.
Publicado: (2025)
por: Song, Zhuoqing, et al.
Publicado: (2025)
A Reduction Algorithm for Markovian Contextual Linear Bandits
por: Buyukkalayci, Kaan, et al.
Publicado: (2026)
por: Buyukkalayci, Kaan, et al.
Publicado: (2026)
A Persistent-State Dataflow Accelerator for Memory-Bound Linear Attention Decode on FPGA
por: Gupta, Neelesh, et al.
Publicado: (2026)
por: Gupta, Neelesh, et al.
Publicado: (2026)
On the Learnability of Offline Model-Based Optimization: A Ranking Perspective
por: Lyu, Shen-Huan, et al.
Publicado: (2026)
por: Lyu, Shen-Huan, et al.
Publicado: (2026)
Low-Rank Learning by Design: the Role of Network Architecture and Activation Linearity in Gradient Rank Collapse
por: Baker, Bradley T., et al.
Publicado: (2024)
por: Baker, Bradley T., et al.
Publicado: (2024)
Cottention: Linear Transformers With Cosine Attention
por: Mongaras, Gabriel, et al.
Publicado: (2024)
por: Mongaras, Gabriel, et al.
Publicado: (2024)
The Rank-Reduced Kalman Filter: Approximate Dynamical-Low-Rank Filtering In High Dimensions
por: Schmidt, Jonathan, et al.
Publicado: (2023)
por: Schmidt, Jonathan, et al.
Publicado: (2023)
Convolutional versus Dense Neural Networks: Comparing the Two Neural Networks Performance in Predicting Building Operational Energy Use Based on the Building Shape
por: Nazari, Farnaz, et al.
Publicado: (2021)
por: Nazari, Farnaz, et al.
Publicado: (2021)
Loki: Low-rank Keys for Efficient Sparse Attention
por: Singhania, Prajwal, et al.
Publicado: (2024)
por: Singhania, Prajwal, et al.
Publicado: (2024)
Exact Linear Attention
por: Ou, Weinuo
Publicado: (2026)
por: Ou, Weinuo
Publicado: (2026)
Kaczmarz Linear Attention
por: Zou, Jiaxuan, et al.
Publicado: (2026)
por: Zou, Jiaxuan, et al.
Publicado: (2026)
Why Softmax Attention Outperforms Linear Attention
por: Deng, Yichuan, et al.
Publicado: (2023)
por: Deng, Yichuan, et al.
Publicado: (2023)
Coupled Query-Key Dynamics for Attention
por: Gahtan, Barak, et al.
Publicado: (2026)
por: Gahtan, Barak, et al.
Publicado: (2026)
Hypernetwork-Driven Low-Rank Adaptation Across Attention Heads
por: Diep, Nghiem T., et al.
Publicado: (2025)
por: Diep, Nghiem T., et al.
Publicado: (2025)
Ejemplares similares
-
The Curious Case of In-Training Compression of State Space Models
por: Chahine, Makram, et al.
Publicado: (2025) -
Oscillatory State-Space Models
por: Rusch, T. Konstantin, et al.
Publicado: (2024) -
Learning to Dissipate Energy in Oscillatory State-Space Models
por: Boyer, Jared, et al.
Publicado: (2025) -
Low-Pass Flow Matching
por: Ruscio, Francesco M., et al.
Publicado: (2026) -
State Rank Dynamics in Linear Attention LLMs
por: Sun, Ao, et al.
Publicado: (2026)