Saved in:
| Main Author: | Karbevski, Marko |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.13381 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers
by: Karbevski, Marko, et al.
Published: (2025)
by: Karbevski, Marko, et al.
Published: (2025)
Can an MLP Absorb Its Own Skip Connection?
by: Mijoski, Antonij, et al.
Published: (2026)
by: Mijoski, Antonij, et al.
Published: (2026)
Exploring Local Explanations of Nonlinear Models Using Animated Linear Projections
by: Spyrison, Nicholas, et al.
Published: (2022)
by: Spyrison, Nicholas, et al.
Published: (2022)
Exact Linear Attention
by: Ou, Weinuo
Published: (2026)
by: Ou, Weinuo
Published: (2026)
Kaczmarz Linear Attention
by: Zou, Jiaxuan, et al.
Published: (2026)
by: Zou, Jiaxuan, et al.
Published: (2026)
Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025)
by: Zuo, Yifei, et al.
Published: (2025)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
by: Hu, Wenjie, et al.
Published: (2025)
by: Hu, Wenjie, et al.
Published: (2025)
On the Identifiability of Nonlinear ICA: Sparsity and Beyond
by: Zheng, Yujia, et al.
Published: (2022)
by: Zheng, Yujia, et al.
Published: (2022)
XicorAttention: Time Series Transformer Using Attention with Nonlinear Correlation
by: Kimura, Daichi, et al.
Published: (2025)
by: Kimura, Daichi, et al.
Published: (2025)
Tree-Sliced Wasserstein Distance with Nonlinear Projection
by: Tran, Thanh, et al.
Published: (2025)
by: Tran, Thanh, et al.
Published: (2025)
Enhancing Linear Attention with Residual Learning
by: Lai, Xunhao, et al.
Published: (2025)
by: Lai, Xunhao, et al.
Published: (2025)
Surprised by Attention: Predictable Query Dynamics for Time Series Anomaly Detection
by: Özer, Kadir-Kaan, et al.
Published: (2026)
by: Özer, Kadir-Kaan, et al.
Published: (2026)
The Laplacian Keyboard: Beyond the Linear Span
by: Chandrasekar, Siddarth, et al.
Published: (2026)
by: Chandrasekar, Siddarth, et al.
Published: (2026)
Adaptive Memory Decay for Log-Linear Attention
by: Amin, Yaxita, et al.
Published: (2026)
by: Amin, Yaxita, et al.
Published: (2026)
KVBuffer: IO-aware Serving for Linear Attention
by: Zou, Longwei, et al.
Published: (2026)
by: Zou, Longwei, et al.
Published: (2026)
State Rank Dynamics in Linear Attention LLMs
by: Sun, Ao, et al.
Published: (2026)
by: Sun, Ao, et al.
Published: (2026)
Linear Attention is Enough in Spatial-Temporal Forecasting
by: Ning, Xinyu
Published: (2024)
by: Ning, Xinyu
Published: (2024)
Linear Attention for Efficient Bidirectional Sequence Modeling
by: Afzal, Arshia, et al.
Published: (2025)
by: Afzal, Arshia, et al.
Published: (2025)
Conservative Contextual Bandits: Beyond Linear Representations
by: Deb, Rohan, et al.
Published: (2024)
by: Deb, Rohan, et al.
Published: (2024)
Linear Projections of Teacher Embeddings for Few-Class Distillation
by: Loo, Noel, et al.
Published: (2024)
by: Loo, Noel, et al.
Published: (2024)
Multi-Item-Query Attention for Stable Sequential Recommendation
by: Xu, Mingshi, et al.
Published: (2025)
by: Xu, Mingshi, et al.
Published: (2025)
GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding
by: Meng, Fanxu
Published: (2026)
by: Meng, Fanxu
Published: (2026)
Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
by: Beck, Maximilian, et al.
Published: (2025)
by: Beck, Maximilian, et al.
Published: (2025)
Nexusformer: Nonlinear Attention Expansion for Stable and Inheritable Transformer Scaling
by: Zhao, Weijie, et al.
Published: (2026)
by: Zhao, Weijie, et al.
Published: (2026)
ZeroS: Zero-Sum Linear Attention for Efficient Transformers
by: Lu, Jiecheng, et al.
Published: (2026)
by: Lu, Jiecheng, et al.
Published: (2026)
Superiority of Multi-Head Attention in In-Context Linear Regression
by: Cui, Yingqian, et al.
Published: (2024)
by: Cui, Yingqian, et al.
Published: (2024)
LinearizeLLM: An Agent-Based Framework for LLM-Driven Exact Linear Reformulation of Nonlinear Optimization Problems
by: Kandora, Paul-Niklas Ken, et al.
Published: (2025)
by: Kandora, Paul-Niklas Ken, et al.
Published: (2025)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
by: Joshi, Sahil, et al.
Published: (2025)
by: Joshi, Sahil, et al.
Published: (2025)
Higher-order Linear Attention
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
From Projection to Prediction: Beyond Logits for Scalable Language Models
by: Dong, Jianbing, et al.
Published: (2025)
by: Dong, Jianbing, et al.
Published: (2025)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
by: Chen, Yingfa, et al.
Published: (2025)
by: Chen, Yingfa, et al.
Published: (2025)
Learning Linear Utility Functions From Pairwise Comparison Queries
by: Ge, Luise, et al.
Published: (2024)
by: Ge, Luise, et al.
Published: (2024)
Beyond Similarity: Temporal Operator Attention for Time Series Analysis
by: Twitty, Jevon, et al.
Published: (2026)
by: Twitty, Jevon, et al.
Published: (2026)
Beyond Additivity: Sparse Isotonic Shapley Regression toward Nonlinear Explainability
by: She, Jialai
Published: (2025)
by: She, Jialai
Published: (2025)
Nonlinear Processing with Linear Optics
by: Yildirim, Mustafa, et al.
Published: (2023)
by: Yildirim, Mustafa, et al.
Published: (2023)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
Learning Advanced Self-Attention for Linear Transformers in the Singular Value Domain
by: Wi, Hyowon, et al.
Published: (2025)
by: Wi, Hyowon, et al.
Published: (2025)
MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map
by: Chou, Yuhong, et al.
Published: (2024)
by: Chou, Yuhong, et al.
Published: (2024)
LNUCB-TA: Linear-nonlinear Hybrid Bandit Learning with Temporal Attention
by: Khosravi, Hamed, et al.
Published: (2025)
by: Khosravi, Hamed, et al.
Published: (2025)
Similar Items
-
Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers
by: Karbevski, Marko, et al.
Published: (2025) -
Can an MLP Absorb Its Own Skip Connection?
by: Mijoski, Antonij, et al.
Published: (2026) -
Exploring Local Explanations of Nonlinear Models Using Animated Linear Projections
by: Spyrison, Nicholas, et al.
Published: (2022) -
Exact Linear Attention
by: Ou, Weinuo
Published: (2026) -
Kaczmarz Linear Attention
by: Zou, Jiaxuan, et al.
Published: (2026)