KVBuffer: IO-aware Serving for Linear Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zou, Longwei, Zhong, Lin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Kaczmarz Linear Attention
von: Zou, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Zou, Jiaxuan, et al.
Veröffentlicht: (2026)
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
von: Guo, Wentao, et al.
Veröffentlicht: (2025)
von: Guo, Wentao, et al.
Veröffentlicht: (2025)
Progressive Sparse Attention: Algorithm and System Co-design for Efficient Attention in LLM Serving
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
von: Zhou, Qihui, et al.
Veröffentlicht: (2025)
Enhancing Linear Attention with Residual Learning
von: Lai, Xunhao, et al.
Veröffentlicht: (2025)
von: Lai, Xunhao, et al.
Veröffentlicht: (2025)
Toward Carbon-Neutral Human AI: Rethinking Data, Computation, and Learning Paradigms for Sustainable Intelligence
von: Santosh, KC, et al.
Veröffentlicht: (2025)
von: Santosh, KC, et al.
Veröffentlicht: (2025)
Enhancing Adversarial Robustness of Deep Neural Networks Through Supervised Contrastive Learning
von: Wang, Longwei, et al.
Veröffentlicht: (2024)
von: Wang, Longwei, et al.
Veröffentlicht: (2024)
Bridging Interpretability and Robustness Using LIME-Guided Model Refinement
von: Nayyem, Navid, et al.
Veröffentlicht: (2024)
von: Nayyem, Navid, et al.
Veröffentlicht: (2024)
Exact Linear Attention
von: Ou, Weinuo
Veröffentlicht: (2026)
von: Ou, Weinuo
Veröffentlicht: (2026)
Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads
von: Song, Chendong, et al.
Veröffentlicht: (2026)
von: Song, Chendong, et al.
Veröffentlicht: (2026)
ATP: Enabling Fast LLM Serving via Attention on Top Principal Keys
von: Niu, Yue, et al.
Veröffentlicht: (2024)
von: Niu, Yue, et al.
Veröffentlicht: (2024)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
von: Zuo, Yifei, et al.
Veröffentlicht: (2025)
von: Zuo, Yifei, et al.
Veröffentlicht: (2025)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
von: Hu, Wenjie, et al.
Veröffentlicht: (2025)
von: Hu, Wenjie, et al.
Veröffentlicht: (2025)
Lipschitz-aware Linearity Grafting for Certified Robustness
von: Han, Yongjin, et al.
Veröffentlicht: (2025)
von: Han, Yongjin, et al.
Veröffentlicht: (2025)
MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map
von: Chou, Yuhong, et al.
Veröffentlicht: (2024)
von: Chou, Yuhong, et al.
Veröffentlicht: (2024)
Adaptive Memory Decay for Log-Linear Attention
von: Amin, Yaxita, et al.
Veröffentlicht: (2026)
von: Amin, Yaxita, et al.
Veröffentlicht: (2026)
State Rank Dynamics in Linear Attention LLMs
von: Sun, Ao, et al.
Veröffentlicht: (2026)
von: Sun, Ao, et al.
Veröffentlicht: (2026)
Linear Attention is Enough in Spatial-Temporal Forecasting
von: Ning, Xinyu
Veröffentlicht: (2024)
von: Ning, Xinyu
Veröffentlicht: (2024)
Linear Attention for Efficient Bidirectional Sequence Modeling
von: Afzal, Arshia, et al.
Veröffentlicht: (2025)
von: Afzal, Arshia, et al.
Veröffentlicht: (2025)
On Efficient Scaling of GNNs via IO-Aware Layers Implementations
von: Fomina, Daria, et al.
Veröffentlicht: (2026)
von: Fomina, Daria, et al.
Veröffentlicht: (2026)
BoA: Attention-aware Post-training Quantization without Backpropagation
von: Kim, Junhan, et al.
Veröffentlicht: (2024)
von: Kim, Junhan, et al.
Veröffentlicht: (2024)
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
Tiled Flash Linear Attention: More Efficient Linear RNN and xLSTM Kernels
von: Beck, Maximilian, et al.
Veröffentlicht: (2025)
von: Beck, Maximilian, et al.
Veröffentlicht: (2025)
Beyond Linearity in Attention Projections: The Case for Nonlinear Queries
von: Karbevski, Marko
Veröffentlicht: (2026)
von: Karbevski, Marko
Veröffentlicht: (2026)
ZeroS: Zero-Sum Linear Attention for Efficient Transformers
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
Superiority of Multi-Head Attention in In-Context Linear Regression
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
E2Former-V2: On-the-Fly Equivariant Attention with Linear Activation Memory
von: Huang, Lin, et al.
Veröffentlicht: (2026)
von: Huang, Lin, et al.
Veröffentlicht: (2026)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
von: Joshi, Sahil, et al.
Veröffentlicht: (2025)
von: Joshi, Sahil, et al.
Veröffentlicht: (2025)
Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
von: Lin, Chaofan, et al.
Veröffentlicht: (2024)
von: Lin, Chaofan, et al.
Veröffentlicht: (2024)
Efficient High-Accuracy PDEs Solver with the Linear Attention Neural Operator
von: Zhong, Ming, et al.
Veröffentlicht: (2025)
von: Zhong, Ming, et al.
Veröffentlicht: (2025)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
von: Ye, Zihao, et al.
Veröffentlicht: (2025)
von: Ye, Zihao, et al.
Veröffentlicht: (2025)
Learning under Quantization for High-Dimensional Linear Regression
von: Zhang, Dechen, et al.
Veröffentlicht: (2025)
von: Zhang, Dechen, et al.
Veröffentlicht: (2025)
Higher-order Linear Attention
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Learning Advanced Self-Attention for Linear Transformers in the Singular Value Domain
von: Wi, Hyowon, et al.
Veröffentlicht: (2025)
von: Wi, Hyowon, et al.
Veröffentlicht: (2025)
LNUCB-TA: Linear-nonlinear Hybrid Bandit Learning with Temporal Attention
von: Khosravi, Hamed, et al.
Veröffentlicht: (2025)
von: Khosravi, Hamed, et al.
Veröffentlicht: (2025)
Revisiting Service Level Objectives and System Level Metrics in Large Language Model Serving
von: Wang, Zhibin, et al.
Veröffentlicht: (2024)
von: Wang, Zhibin, et al.
Veröffentlicht: (2024)
When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
von: You, Haoran, et al.
Veröffentlicht: (2024)
von: You, Haoran, et al.
Veröffentlicht: (2024)
Distance-aware Attention Reshaping: Enhance Generalization of Neural Solver for Large-scale Vehicle Routing Problems
von: Wang, Yang, et al.
Veröffentlicht: (2024)
von: Wang, Yang, et al.
Veröffentlicht: (2024)
TA-RNN: an Attention-based Time-aware Recurrent Neural Network Architecture for Electronic Health Records
von: Olaimat, Mohammad Al, et al.
Veröffentlicht: (2024)
von: Olaimat, Mohammad Al, et al.
Veröffentlicht: (2024)
MTServe: Efficient Serving for Generative Recommendation Models with Hierarchical Caches
von: Wang, Xin, et al.
Veröffentlicht: (2026)
von: Wang, Xin, et al.
Veröffentlicht: (2026)
Scaling Laws for Precision in High-Dimensional Linear Regression
von: Zhang, Dechen, et al.
Veröffentlicht: (2026)
von: Zhang, Dechen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Kaczmarz Linear Attention
von: Zou, Jiaxuan, et al.
Veröffentlicht: (2026) -
SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
von: Guo, Wentao, et al.
Veröffentlicht: (2025) -
Progressive Sparse Attention: Algorithm and System Co-design for Efficient Attention in LLM Serving
von: Zhou, Qihui, et al.
Veröffentlicht: (2025) -
Enhancing Linear Attention with Residual Learning
von: Lai, Xunhao, et al.
Veröffentlicht: (2025) -
Toward Carbon-Neutral Human AI: Rethinking Data, Computation, and Learning Paradigms for Sustainable Intelligence
von: Santosh, KC, et al.
Veröffentlicht: (2025)