One Pass Streaming Algorithm for Super Long Token Attention Approximation in Sublinear Space
Fuente:
arXiv
Saved in:
| Main Authors: | Addanki, Raghav, Li, Chenyang, Song, Zhao, Yang, Chiwun |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
Towards Infinite-Long Prefix in Transformer
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
Unifying Learning Dynamics and Generalization in Transformers Scaling Law
by: Yang, Chiwun
Published: (2025)
by: Yang, Chiwun
Published: (2025)
Unlocking the Theory Behind Scaling 1-Bit Neural Networks
by: Daliri, Majid, et al.
Published: (2024)
by: Daliri, Majid, et al.
Published: (2024)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
On Fine-Grained I/O Complexity of Attention Backward Passes
by: Li, Xiaoyu, et al.
Published: (2024)
by: Li, Xiaoyu, et al.
Published: (2024)
Sublinear Time Quantum Algorithm for Attention Approximation
by: Song, Zhao, et al.
Published: (2026)
by: Song, Zhao, et al.
Published: (2026)
Introducing Super RAGs in Mistral 8x7B-v1
by: Thakur, Ayush, et al.
Published: (2024)
by: Thakur, Ayush, et al.
Published: (2024)
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning
by: Zarch, Hossein Entezari, et al.
Published: (2025)
by: Zarch, Hossein Entezari, et al.
Published: (2025)
Theoretical Foundation of Flow-Based Time Series Generation: Provable Approximation, Generalization, and Efficiency
by: Long, Jiangxuan, et al.
Published: (2025)
by: Long, Jiangxuan, et al.
Published: (2025)
Rethinking Token Reduction for State Space Models
by: Zhan, Zheng, et al.
Published: (2024)
by: Zhan, Zheng, et al.
Published: (2024)
One Token to Fool LLM-as-a-Judge
by: Zhao, Yulai, et al.
Published: (2025)
by: Zhao, Yulai, et al.
Published: (2025)
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
by: Chen, Yingfa, et al.
Published: (2025)
by: Chen, Yingfa, et al.
Published: (2025)
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization
by: Song, Guanghui, et al.
Published: (2025)
by: Song, Guanghui, et al.
Published: (2025)
One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning
by: Goru, Ritesh, et al.
Published: (2025)
by: Goru, Ritesh, et al.
Published: (2025)
Token Distillation: Attention-aware Input Embeddings For New Tokens
by: Dobler, Konstantin, et al.
Published: (2025)
by: Dobler, Konstantin, et al.
Published: (2025)
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
by: Tang, Hongyin, et al.
Published: (2024)
by: Tang, Hongyin, et al.
Published: (2024)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
by: Huang, Yuxiang, et al.
Published: (2025)
by: Huang, Yuxiang, et al.
Published: (2025)
Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
When Every Token Counts: Optimal Segmentation for Low-Resource Language Models
by: Raj, Bharath, et al.
Published: (2024)
by: Raj, Bharath, et al.
Published: (2024)
Softmax Attention with Constant Cost per Token
by: Heinsen, Franz A.
Published: (2024)
by: Heinsen, Franz A.
Published: (2024)
TCUQ: Single-Pass Uncertainty Quantification from Temporal Consistency with Streaming Conformal Calibration for TinyML
by: Lamaakal, Ismail, et al.
Published: (2025)
by: Lamaakal, Ismail, et al.
Published: (2025)
Why Softmax Attention Outperforms Linear Attention
by: Deng, Yichuan, et al.
Published: (2023)
by: Deng, Yichuan, et al.
Published: (2023)
STS: Efficient Sparse Attention with Speculative Token Sparsity
by: Xu, Ceyu, et al.
Published: (2026)
by: Xu, Ceyu, et al.
Published: (2026)
Attention with Trained Embeddings Provably Selects Important Tokens
by: Wu, Diyuan, et al.
Published: (2025)
by: Wu, Diyuan, et al.
Published: (2025)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
by: He, Mutian, et al.
Published: (2025)
by: He, Mutian, et al.
Published: (2025)
Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models
by: Deng, Difan, et al.
Published: (2026)
by: Deng, Difan, et al.
Published: (2026)
SuperBPE: Space Travel for Language Models
by: Liu, Alisa, et al.
Published: (2025)
by: Liu, Alisa, et al.
Published: (2025)
Multipole Attention for Efficient Long Context Reasoning
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
Aligning Multiple Knowledge Graphs in a Single Pass
by: Yang, Yaming, et al.
Published: (2024)
by: Yang, Yaming, et al.
Published: (2024)
Nectar: Neural Estimation of Cached-Token Attention via Regression
by: Monteiro, João, et al.
Published: (2026)
by: Monteiro, João, et al.
Published: (2026)
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
by: Hu, Junhao, et al.
Published: (2026)
by: Hu, Junhao, et al.
Published: (2026)
Dist2ill: Distributional Distillation for One-Pass Uncertainty Estimation in Large Language Models
by: Zhao, Yicong, et al.
Published: (2025)
by: Zhao, Yicong, et al.
Published: (2025)
LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
by: Zhang, Haoyue, et al.
Published: (2025)
by: Zhang, Haoyue, et al.
Published: (2025)
Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
by: Su, Guinan, et al.
Published: (2026)
by: Su, Guinan, et al.
Published: (2026)
Learning to Explain: Supervised Token Attribution from Transformer Attention Patterns
by: Mihaila, George
Published: (2026)
by: Mihaila, George
Published: (2026)
QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm
by: Zhou, Qirui, et al.
Published: (2025)
by: Zhou, Qirui, et al.
Published: (2025)
Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention
by: Zhan, Zhihao, et al.
Published: (2025)
by: Zhan, Zhihao, et al.
Published: (2025)
Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation
by: Phan, Buu, et al.
Published: (2025)
by: Phan, Buu, et al.
Published: (2025)
Similar Items
-
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024) -
Towards Infinite-Long Prefix in Transformer
by: Liang, Yingyu, et al.
Published: (2024) -
Unifying Learning Dynamics and Generalization in Transformers Scaling Law
by: Yang, Chiwun
Published: (2025) -
Unlocking the Theory Behind Scaling 1-Bit Neural Networks
by: Daliri, Majid, et al.
Published: (2024) -
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)