NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Tianyi, Yi, Jonah Wonkyu, Yao, Bowen, Xu, Zhaozhuo, Shrivastava, Anshumali |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
di: Le, Hoang Anh Duy, et al.
Pubblicazione: (2026)
di: Le, Hoang Anh Duy, et al.
Pubblicazione: (2026)
KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
di: Yang, Zeyu, et al.
Pubblicazione: (2025)
di: Yang, Zeyu, et al.
Pubblicazione: (2025)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
di: Joshi, Sahil, et al.
Pubblicazione: (2025)
di: Joshi, Sahil, et al.
Pubblicazione: (2025)
Star Attention: Efficient LLM Inference over Long Sequences
di: Acharya, Shantanu, et al.
Pubblicazione: (2024)
di: Acharya, Shantanu, et al.
Pubblicazione: (2024)
Highly Optimized Kernels and Fine-Grained Codebooks for LLM Inference on Arm CPUs
di: Gope, Dibakar, et al.
Pubblicazione: (2024)
di: Gope, Dibakar, et al.
Pubblicazione: (2024)
REFRAG: Rethinking RAG based Decoding
di: Lin, Xiaoqiang, et al.
Pubblicazione: (2025)
di: Lin, Xiaoqiang, et al.
Pubblicazione: (2025)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
di: Zhu, Qianchao, et al.
Pubblicazione: (2024)
di: Zhu, Qianchao, et al.
Pubblicazione: (2024)
Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM Adaptation
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
di: Lv, Xingtai, et al.
Pubblicazione: (2024)
di: Lv, Xingtai, et al.
Pubblicazione: (2024)
Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
di: Sharma, Agniv, et al.
Pubblicazione: (2024)
di: Sharma, Agniv, et al.
Pubblicazione: (2024)
AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs
di: S, Santhosh G, et al.
Pubblicazione: (2025)
di: S, Santhosh G, et al.
Pubblicazione: (2025)
Block-Attention for Efficient Prefilling
di: Ma, Dongyang, et al.
Pubblicazione: (2024)
di: Ma, Dongyang, et al.
Pubblicazione: (2024)
iMAD: Intelligent Multi-Agent Debate for Efficient and Accurate LLM Inference
di: Fan, Wei, et al.
Pubblicazione: (2025)
di: Fan, Wei, et al.
Pubblicazione: (2025)
Sliding Window Attention Training for Efficient Large Language Models
di: Fu, Zichuan, et al.
Pubblicazione: (2025)
di: Fu, Zichuan, et al.
Pubblicazione: (2025)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
di: Huang, Yuxiang, et al.
Pubblicazione: (2026)
di: Huang, Yuxiang, et al.
Pubblicazione: (2026)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
di: Wang, Guangtao, et al.
Pubblicazione: (2025)
di: Wang, Guangtao, et al.
Pubblicazione: (2025)
Native Hybrid Attention for Efficient Sequence Modeling
di: Du, Jusen, et al.
Pubblicazione: (2025)
di: Du, Jusen, et al.
Pubblicazione: (2025)
Round Attention: A Novel Round-Level Attention Mechanism to Accelerate LLM Inference
di: Tang, Yaohua, et al.
Pubblicazione: (2025)
di: Tang, Yaohua, et al.
Pubblicazione: (2025)
Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps
di: Waldendorf, Jonas, et al.
Pubblicazione: (2026)
di: Waldendorf, Jonas, et al.
Pubblicazione: (2026)
The Map of Misbelief: Tracing Intrinsic and Extrinsic Hallucinations Through Attention Patterns
di: Hajji, Elyes, et al.
Pubblicazione: (2025)
di: Hajji, Elyes, et al.
Pubblicazione: (2025)
FROST: Filtering Reasoning Outliers with Attention for Efficient Reasoning
di: Luo, Haozheng, et al.
Pubblicazione: (2026)
di: Luo, Haozheng, et al.
Pubblicazione: (2026)
Beyond KV Caching: Shared Attention for Efficient LLMs
di: Liao, Bingli, et al.
Pubblicazione: (2024)
di: Liao, Bingli, et al.
Pubblicazione: (2024)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
di: Liu, Guangda, et al.
Pubblicazione: (2025)
di: Liu, Guangda, et al.
Pubblicazione: (2025)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
di: MiniCPM Team, et al.
Pubblicazione: (2026)
di: MiniCPM Team, et al.
Pubblicazione: (2026)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
di: Agrawal, Sudhanshu, et al.
Pubblicazione: (2025)
di: Agrawal, Sudhanshu, et al.
Pubblicazione: (2025)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
di: Fu, Tianyu, et al.
Pubblicazione: (2024)
di: Fu, Tianyu, et al.
Pubblicazione: (2024)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
di: Yang, Shang, et al.
Pubblicazione: (2025)
di: Yang, Shang, et al.
Pubblicazione: (2025)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
di: Ling Team, et al.
Pubblicazione: (2025)
di: Ling Team, et al.
Pubblicazione: (2025)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
di: Deng, Yichuan, et al.
Pubblicazione: (2024)
di: Deng, Yichuan, et al.
Pubblicazione: (2024)
Superintelligent Retrieval Agent: The Next Frontier of Information Retrieval
di: Yang, Zeyu, et al.
Pubblicazione: (2026)
di: Yang, Zeyu, et al.
Pubblicazione: (2026)
Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries
di: Hicke, Rebecca M. M., et al.
Pubblicazione: (2026)
di: Hicke, Rebecca M. M., et al.
Pubblicazione: (2026)
Scaling Stick-Breaking Attention: An Efficient Implementation and In-depth Study
di: Tan, Shawn, et al.
Pubblicazione: (2024)
di: Tan, Shawn, et al.
Pubblicazione: (2024)
MOM: Memory-Efficient Offloaded Mini-Sequence Inference for Long Context Language Models
di: Zhang, Junyang, et al.
Pubblicazione: (2025)
di: Zhang, Junyang, et al.
Pubblicazione: (2025)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
di: Fu, Zichuan, et al.
Pubblicazione: (2026)
di: Fu, Zichuan, et al.
Pubblicazione: (2026)
DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial Attention
di: Lee, Younjoo, et al.
Pubblicazione: (2026)
di: Lee, Younjoo, et al.
Pubblicazione: (2026)
Tensor Product Attention Is All You Need
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
di: Zhao, Weilin, et al.
Pubblicazione: (2025)
di: Zhao, Weilin, et al.
Pubblicazione: (2025)
DEL-ToM: Inference-Time Scaling for Theory-of-Mind Reasoning via Dynamic Epistemic Logic
di: Wu, Yuheng, et al.
Pubblicazione: (2025)
di: Wu, Yuheng, et al.
Pubblicazione: (2025)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
di: Yang, Lijie, et al.
Pubblicazione: (2024)
di: Yang, Lijie, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
di: Le, Hoang Anh Duy, et al.
Pubblicazione: (2026) -
KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization
di: Zhang, Tianyi, et al.
Pubblicazione: (2024) -
To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
di: Yang, Zeyu, et al.
Pubblicazione: (2025) -
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
di: Joshi, Sahil, et al.
Pubblicazione: (2025) -
Star Attention: Efficient LLM Inference over Long Sequences
di: Acharya, Shantanu, et al.
Pubblicazione: (2024)