Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Norgren, Victor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stateful Inference for Low-Latency Multi-Agent Tool Calling
von: Norgren, Victor
Veröffentlicht: (2026)
von: Norgren, Victor
Veröffentlicht: (2026)
The Residual Stream Is All You Need: On the Redundancy of the KV Cache in Transformer Inference
von: Qasim, Kaleem Ullah, et al.
Veröffentlicht: (2026)
von: Qasim, Kaleem Ullah, et al.
Veröffentlicht: (2026)
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
von: Tyukin, Georgy, et al.
Veröffentlicht: (2024)
von: Tyukin, Georgy, et al.
Veröffentlicht: (2024)
Attention is All You Need Until You Need Retention
von: Yaslioglu, M. Murat
Veröffentlicht: (2025)
von: Yaslioglu, M. Murat
Veröffentlicht: (2025)
Some Attention is All You Need for Retrieval
von: Michalak, Felix, et al.
Veröffentlicht: (2025)
von: Michalak, Felix, et al.
Veröffentlicht: (2025)
Attention Is Not All You Need: The Importance of Feedforward Networks in Transformer Models
von: Gerber, Isaac
Veröffentlicht: (2025)
von: Gerber, Isaac
Veröffentlicht: (2025)
Element-wise Attention Is All You Need
von: Feng, Guoxin
Veröffentlicht: (2025)
von: Feng, Guoxin
Veröffentlicht: (2025)
Attention Smoothing Is All You Need For Unlearning
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026)
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026)
Tensor Product Attention Is All You Need
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Accuracy is Not All You Need
von: Dutta, Abhinav, et al.
Veröffentlicht: (2024)
von: Dutta, Abhinav, et al.
Veröffentlicht: (2024)
TransMLA: Multi-Head Latent Attention Is All You Need
von: Meng, Fanxu, et al.
Veröffentlicht: (2025)
von: Meng, Fanxu, et al.
Veröffentlicht: (2025)
Attention is All You Need to Optimize Wind Farm Operations and Maintenance
von: Kazemian, Iman, et al.
Veröffentlicht: (2024)
von: Kazemian, Iman, et al.
Veröffentlicht: (2024)
Attention Is All You Need for KV Cache in Diffusion LLMs
von: Nguyen-Tri, Quan, et al.
Veröffentlicht: (2025)
von: Nguyen-Tri, Quan, et al.
Veröffentlicht: (2025)
Context is All You Need
von: Delanois, Jean Erik, et al.
Veröffentlicht: (2026)
von: Delanois, Jean Erik, et al.
Veröffentlicht: (2026)
What Matters in Transformers? Not All Attention is Needed
von: He, Shwai, et al.
Veröffentlicht: (2024)
von: He, Shwai, et al.
Veröffentlicht: (2024)
Forget Attention: Importance-Aware Attention Is All You Need
von: Shin, Soohyeong, et al.
Veröffentlicht: (2026)
von: Shin, Soohyeong, et al.
Veröffentlicht: (2026)
Top-$nσ$: Not All Logits Are You Need
von: Tang, Chenxia, et al.
Veröffentlicht: (2024)
von: Tang, Chenxia, et al.
Veröffentlicht: (2024)
Half Search Space is All You Need
von: Rumiantsev, Pavel, et al.
Veröffentlicht: (2025)
von: Rumiantsev, Pavel, et al.
Veröffentlicht: (2025)
Context-Selective State Space Models: Feedback is All You Need
von: Zattra, Riccardo, et al.
Veröffentlicht: (2025)
von: Zattra, Riccardo, et al.
Veröffentlicht: (2025)
Efficient Deep Learning Board: Training Feedback Is Not All You Need
von: Gong, Lina, et al.
Veröffentlicht: (2024)
von: Gong, Lina, et al.
Veröffentlicht: (2024)
Attentional Graph Neural Network Is All You Need for Robust Massive Network Localization
von: Yan, Wenzhong, et al.
Veröffentlicht: (2023)
von: Yan, Wenzhong, et al.
Veröffentlicht: (2023)
Exploitation Is All You Need... for Exploration
von: Rentschler, Micah, et al.
Veröffentlicht: (2025)
von: Rentschler, Micah, et al.
Veröffentlicht: (2025)
Multistep Inverse Is Not All You Need
von: Levine, Alexander, et al.
Veröffentlicht: (2024)
von: Levine, Alexander, et al.
Veröffentlicht: (2024)
Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers
von: Karbevski, Marko, et al.
Veröffentlicht: (2025)
von: Karbevski, Marko, et al.
Veröffentlicht: (2025)
Embedding Is (Almost) All You Need: Retrieval-Augmented Inference for Generalizable Genomic Prediction Tasks
von: Datta, Nirjhor, et al.
Veröffentlicht: (2025)
von: Datta, Nirjhor, et al.
Veröffentlicht: (2025)
MoE Lens -- An Expert Is All You Need
von: Chaudhari, Marmik, et al.
Veröffentlicht: (2026)
von: Chaudhari, Marmik, et al.
Veröffentlicht: (2026)
Support is All You Need for Certified VAE Training
von: Xu, Changming, et al.
Veröffentlicht: (2025)
von: Xu, Changming, et al.
Veröffentlicht: (2025)
Fusion or Confusion? Multimodal Complexity Is Not All You Need
von: Rheude, Tillmann, et al.
Veröffentlicht: (2025)
von: Rheude, Tillmann, et al.
Veröffentlicht: (2025)
Realizable Learning is All You Need
von: Hopkins, Max, et al.
Veröffentlicht: (2021)
von: Hopkins, Max, et al.
Veröffentlicht: (2021)
You Need Better Attention Priors
von: Litman, Elon, et al.
Veröffentlicht: (2026)
von: Litman, Elon, et al.
Veröffentlicht: (2026)
Cooperation Is All You Need
von: Adeel, Ahsan, et al.
Veröffentlicht: (2023)
von: Adeel, Ahsan, et al.
Veröffentlicht: (2023)
Catch-Up Distillation: You Only Need to Train Once for Accelerating Sampling
von: Shao, Shitong, et al.
Veröffentlicht: (2023)
von: Shao, Shitong, et al.
Veröffentlicht: (2023)
You Only Accept Samples Once: Fast, Self-Correcting Stochastic Variational Inference
von: Dayta, Dominic B.
Veröffentlicht: (2024)
von: Dayta, Dominic B.
Veröffentlicht: (2024)
You Only Debias Once: Towards Flexible Accuracy-Fairness Trade-offs at Inference Time
von: Han, Xiaotian, et al.
Veröffentlicht: (2025)
von: Han, Xiaotian, et al.
Veröffentlicht: (2025)
Attention Is Not What You Need
von: Chong, Zhang
Veröffentlicht: (2025)
von: Chong, Zhang
Veröffentlicht: (2025)
Attention Isn't All You Need for Emotion Recognition:Domain Features Outperform Transformers on the EAV Dataset
von: Guragain, Anmol
Veröffentlicht: (2026)
von: Guragain, Anmol
Veröffentlicht: (2026)
RecurFormer: Not All Transformer Heads Need Self-Attention
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
Positional Knowledge is All You Need: Position-induced Transformer (PiT) for Operator Learning
von: Chen, Junfeng, et al.
Veröffentlicht: (2024)
von: Chen, Junfeng, et al.
Veröffentlicht: (2024)
All You Need Is Synthetic Task Augmentation
von: Godin, Guillaume
Veröffentlicht: (2025)
von: Godin, Guillaume
Veröffentlicht: (2025)
CAMformer: Associative Memory is All You Need
von: Molom-Ochir, Tergel, et al.
Veröffentlicht: (2025)
von: Molom-Ochir, Tergel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Stateful Inference for Low-Latency Multi-Agent Tool Calling
von: Norgren, Victor
Veröffentlicht: (2026) -
The Residual Stream Is All You Need: On the Redundancy of the KV Cache in Transformer Inference
von: Qasim, Kaleem Ullah, et al.
Veröffentlicht: (2026) -
Attention Is All You Need But You Don't Need All Of It For Inference of Large Language Models
von: Tyukin, Georgy, et al.
Veröffentlicht: (2024) -
Attention is All You Need Until You Need Retention
von: Yaslioglu, M. Murat
Veröffentlicht: (2025) -
Some Attention is All You Need for Retrieval
von: Michalak, Felix, et al.
Veröffentlicht: (2025)