The Key to State Reduction in Linear Attention: A Rank-based Perspective
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Nazari, Philipp, Rusch, T. Konstantin |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Curious Case of In-Training Compression of State Space Models
par: Chahine, Makram, et autres
Publié: (2025)
par: Chahine, Makram, et autres
Publié: (2025)
Oscillatory State-Space Models
par: Rusch, T. Konstantin, et autres
Publié: (2024)
par: Rusch, T. Konstantin, et autres
Publié: (2024)
Learning to Dissipate Energy in Oscillatory State-Space Models
par: Boyer, Jared, et autres
Publié: (2025)
par: Boyer, Jared, et autres
Publié: (2025)
Low-Pass Flow Matching
par: Ruscio, Francesco M., et autres
Publié: (2026)
par: Ruscio, Francesco M., et autres
Publié: (2026)
State Rank Dynamics in Linear Attention LLMs
par: Sun, Ao, et autres
Publié: (2026)
par: Sun, Ao, et autres
Publié: (2026)
Low-Rank Key Value Attention
par: O'Neill, James, et autres
Publié: (2026)
par: O'Neill, James, et autres
Publié: (2026)
Quantifying Memory Use in Reinforcement Learning with Temporal Range
par: Lafuente-Mercado, Rodney, et autres
Publié: (2025)
par: Lafuente-Mercado, Rodney, et autres
Publié: (2025)
Low Stein Discrepancy via Message-Passing Monte Carlo
par: Kirk, Nathan, et autres
Publié: (2025)
par: Kirk, Nathan, et autres
Publié: (2025)
Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
par: Boursier, Etienne, et autres
Publié: (2025)
par: Boursier, Etienne, et autres
Publié: (2025)
Relaxed Equivariance via Multitask Learning
par: Elhag, Ahmed A., et autres
Publié: (2024)
par: Elhag, Ahmed A., et autres
Publié: (2024)
LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
par: Shahbazi, Ashkan, et autres
Publié: (2025)
par: Shahbazi, Ashkan, et autres
Publié: (2025)
LoLA: Low-Rank Linear Attention With Sparse Caching
par: McDermott, Luke, et autres
Publié: (2025)
par: McDermott, Luke, et autres
Publié: (2025)
Rank Reduction Autoencoders
par: Mounayer, Jad, et autres
Publié: (2024)
par: Mounayer, Jad, et autres
Publié: (2024)
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction
par: Wu, Ziyang, et autres
Publié: (2024)
par: Wu, Ziyang, et autres
Publié: (2024)
Variational Rank Reduction Autoencoders
par: Mounayer, Jad, et autres
Publié: (2025)
par: Mounayer, Jad, et autres
Publié: (2025)
Power-based Partial Attention: Bridging Linear-Complexity and Full Attention
par: Huang, Yufeng
Publié: (2026)
par: Huang, Yufeng
Publié: (2026)
On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention
par: Ro, Yeonju, et autres
Publié: (2025)
par: Ro, Yeonju, et autres
Publié: (2025)
Message-Passing Monte Carlo: Generating low-discrepancy point sets via Graph Neural Networks
par: Rusch, T. Konstantin, et autres
Publié: (2024)
par: Rusch, T. Konstantin, et autres
Publié: (2024)
Scaling Linear Attention with Sparse State Expansion
par: Pan, Yuqi, et autres
Publié: (2025)
par: Pan, Yuqi, et autres
Publié: (2025)
On the Benefits of Rank in Attention Layers
par: Amsel, Noah, et autres
Publié: (2024)
par: Amsel, Noah, et autres
Publié: (2024)
Neural Low-Discrepancy Sequences
par: Van Huffel, Michael Etienne, et autres
Publié: (2025)
par: Van Huffel, Michael Etienne, et autres
Publié: (2025)
Multi-Head Low-Rank Attention
par: Liu, Songtao, et autres
Publié: (2026)
par: Liu, Songtao, et autres
Publié: (2026)
Log-Linear Attention
par: Guo, Han, et autres
Publié: (2025)
par: Guo, Han, et autres
Publié: (2025)
Kalman Linear Attention: Parallel Bayesian Filtering For Efficient Language Modelling and State Tracking
par: Shaj, Vaisakh, et autres
Publié: (2026)
par: Shaj, Vaisakh, et autres
Publié: (2026)
How does over-squashing affect the power of GNNs?
par: Di Giovanni, Francesco, et autres
Publié: (2023)
par: Di Giovanni, Francesco, et autres
Publié: (2023)
Low-Rank Tensor Decompositions for the Theory of Neural Networks
par: Borsoi, Ricardo, et autres
Publié: (2025)
par: Borsoi, Ricardo, et autres
Publié: (2025)
Causal Attention with Lookahead Keys
par: Song, Zhuoqing, et autres
Publié: (2025)
par: Song, Zhuoqing, et autres
Publié: (2025)
A Reduction Algorithm for Markovian Contextual Linear Bandits
par: Buyukkalayci, Kaan, et autres
Publié: (2026)
par: Buyukkalayci, Kaan, et autres
Publié: (2026)
A Persistent-State Dataflow Accelerator for Memory-Bound Linear Attention Decode on FPGA
par: Gupta, Neelesh, et autres
Publié: (2026)
par: Gupta, Neelesh, et autres
Publié: (2026)
On the Learnability of Offline Model-Based Optimization: A Ranking Perspective
par: Lyu, Shen-Huan, et autres
Publié: (2026)
par: Lyu, Shen-Huan, et autres
Publié: (2026)
Low-Rank Learning by Design: the Role of Network Architecture and Activation Linearity in Gradient Rank Collapse
par: Baker, Bradley T., et autres
Publié: (2024)
par: Baker, Bradley T., et autres
Publié: (2024)
Cottention: Linear Transformers With Cosine Attention
par: Mongaras, Gabriel, et autres
Publié: (2024)
par: Mongaras, Gabriel, et autres
Publié: (2024)
The Rank-Reduced Kalman Filter: Approximate Dynamical-Low-Rank Filtering In High Dimensions
par: Schmidt, Jonathan, et autres
Publié: (2023)
par: Schmidt, Jonathan, et autres
Publié: (2023)
Convolutional versus Dense Neural Networks: Comparing the Two Neural Networks Performance in Predicting Building Operational Energy Use Based on the Building Shape
par: Nazari, Farnaz, et autres
Publié: (2021)
par: Nazari, Farnaz, et autres
Publié: (2021)
Loki: Low-rank Keys for Efficient Sparse Attention
par: Singhania, Prajwal, et autres
Publié: (2024)
par: Singhania, Prajwal, et autres
Publié: (2024)
Exact Linear Attention
par: Ou, Weinuo
Publié: (2026)
par: Ou, Weinuo
Publié: (2026)
Kaczmarz Linear Attention
par: Zou, Jiaxuan, et autres
Publié: (2026)
par: Zou, Jiaxuan, et autres
Publié: (2026)
Why Softmax Attention Outperforms Linear Attention
par: Deng, Yichuan, et autres
Publié: (2023)
par: Deng, Yichuan, et autres
Publié: (2023)
Coupled Query-Key Dynamics for Attention
par: Gahtan, Barak, et autres
Publié: (2026)
par: Gahtan, Barak, et autres
Publié: (2026)
Hypernetwork-Driven Low-Rank Adaptation Across Attention Heads
par: Diep, Nghiem T., et autres
Publié: (2025)
par: Diep, Nghiem T., et autres
Publié: (2025)
Documents similaires
-
The Curious Case of In-Training Compression of State Space Models
par: Chahine, Makram, et autres
Publié: (2025) -
Oscillatory State-Space Models
par: Rusch, T. Konstantin, et autres
Publié: (2024) -
Learning to Dissipate Energy in Oscillatory State-Space Models
par: Boyer, Jared, et autres
Publié: (2025) -
Low-Pass Flow Matching
par: Ruscio, Francesco M., et autres
Publié: (2026) -
State Rank Dynamics in Linear Attention LLMs
par: Sun, Ao, et autres
Publié: (2026)