Squeezed Attention: Accelerating Long Context Length LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Hooper, Coleman, Kim, Sehoon, Mohammadzadeh, Hiva, Maheswaran, Monishwaran, Zhao, Sebastian, Paik, June, Mahoney, Michael W., Keutzer, Kurt, Gholami, Amir |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
by: Hooper, Coleman, et al.
Published: (2024)
by: Hooper, Coleman, et al.
Published: (2024)
SPEED: Speculative Pipelined Execution for Efficient Decoding
by: Hooper, Coleman, et al.
Published: (2023)
by: Hooper, Coleman, et al.
Published: (2023)
Multipole Attention for Efficient Long Context Reasoning
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
ETS: Efficient Tree Search for Inference-Time Scaling
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
SqueezeLLM: Dense-and-Sparse Quantization
by: Kim, Sehoon, et al.
Published: (2023)
by: Kim, Sehoon, et al.
Published: (2023)
AI and Memory Wall
by: Gholami, Amir, et al.
Published: (2024)
by: Gholami, Amir, et al.
Published: (2024)
Characterizing Prompt Compression Methods for Long Context Inference
by: Jha, Siddharth, et al.
Published: (2024)
by: Jha, Siddharth, et al.
Published: (2024)
Learned Best-Effort LLM Serving
by: Jha, Siddharth, et al.
Published: (2024)
by: Jha, Siddharth, et al.
Published: (2024)
Residual Context Diffusion Language Models
by: Hu, Yuezhou, et al.
Published: (2026)
by: Hu, Yuezhou, et al.
Published: (2026)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
by: Maheswaran, Monishwaran, et al.
Published: (2025)
by: Maheswaran, Monishwaran, et al.
Published: (2025)
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
by: Tiwari, Rishabh, et al.
Published: (2026)
by: Tiwari, Rishabh, et al.
Published: (2026)
An LLM Compiler for Parallel Function Calling
by: Kim, Sehoon, et al.
Published: (2023)
by: Kim, Sehoon, et al.
Published: (2023)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
by: Tiwari, Rishabh, et al.
Published: (2025)
by: Tiwari, Rishabh, et al.
Published: (2025)
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
by: Tomar, Aditya, et al.
Published: (2025)
by: Tomar, Aditya, et al.
Published: (2025)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
by: Kim, Minseo, et al.
Published: (2025)
by: Kim, Minseo, et al.
Published: (2025)
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement
by: Lee, Nicholas, et al.
Published: (2024)
by: Lee, Nicholas, et al.
Published: (2024)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
TinyAgent: Function Calling at the Edge
by: Erdogan, Lutfi Eren, et al.
Published: (2024)
by: Erdogan, Lutfi Eren, et al.
Published: (2024)
CDLM: Consistency Diffusion Language Models For Faster Sampling
by: Kim, Minseo, et al.
Published: (2025)
by: Kim, Minseo, et al.
Published: (2025)
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
by: Erdogan, Lutfi Eren, et al.
Published: (2025)
by: Erdogan, Lutfi Eren, et al.
Published: (2025)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
by: Xi, Haocheng, et al.
Published: (2026)
by: Xi, Haocheng, et al.
Published: (2026)
Squeeze Evolve: Unified Multi-Model Orchestration for Verifier-Free Evolution
by: Maheswaran, Monishwaran, et al.
Published: (2026)
by: Maheswaran, Monishwaran, et al.
Published: (2026)
Efficient and Scalable Estimation of Tool Representations in Vector Space
by: Moon, Suhong, et al.
Published: (2024)
by: Moon, Suhong, et al.
Published: (2024)
Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
by: Hooper, Coleman, et al.
Published: (2026)
by: Hooper, Coleman, et al.
Published: (2026)
Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior
by: Subramanian, Shashank, et al.
Published: (2023)
by: Subramanian, Shashank, et al.
Published: (2023)
TASER: Translation Assessment via Systematic Evaluation and Reasoning
by: Maheswaran, Monishwaran, et al.
Published: (2025)
by: Maheswaran, Monishwaran, et al.
Published: (2025)
Agentic Test-Time Scaling for WebAgents
by: Lee, Nicholas, et al.
Published: (2026)
by: Lee, Nicholas, et al.
Published: (2026)
SciML Agents: Write the Solver, Not the Solution
by: Gaonkar, Saarth, et al.
Published: (2025)
by: Gaonkar, Saarth, et al.
Published: (2025)
MITRA: A Large-Scale Parallel Corpus and Multilingual Pretrained Language Model for Machine Translation and Semantic Retrieval for Pāli, Sanskrit, Buddhist Chinese, and Tibetan
by: Nehrdich, Sebastian, et al.
Published: (2026)
by: Nehrdich, Sebastian, et al.
Published: (2026)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
by: Liu, Di, et al.
Published: (2024)
by: Liu, Di, et al.
Published: (2024)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
by: Zhu, Qianchao, et al.
Published: (2024)
by: Zhu, Qianchao, et al.
Published: (2024)
VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference
by: Liu, Anmin, et al.
Published: (2026)
by: Liu, Anmin, et al.
Published: (2026)
One Model is All You Need: ByT5-Sanskrit, a Unified Model for Sanskrit NLP Tasks
by: Nehrdich, Sebastian, et al.
Published: (2024)
by: Nehrdich, Sebastian, et al.
Published: (2024)
A Study on Context Length and Efficient Transformers for Biomedical Image Analysis
by: Hooper, Sarah M., et al.
Published: (2024)
by: Hooper, Sarah M., et al.
Published: (2024)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
by: Gu, Zhuohan, et al.
Published: (2024)
by: Gu, Zhuohan, et al.
Published: (2024)
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
by: Deng, Weishu, et al.
Published: (2025)
by: Deng, Weishu, et al.
Published: (2025)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
by: Zhou, Ruijie, et al.
Published: (2026)
by: Zhou, Ruijie, et al.
Published: (2026)
ODE$_t$(ODE$_l$): Shortcutting the Time and the Length in Diffusion and Flow Models for Faster Sampling
by: Gudovskiy, Denis, et al.
Published: (2025)
by: Gudovskiy, Denis, et al.
Published: (2025)
SparseLoRA: Accelerating LLM Fine-Tuning with Contextual Sparsity
by: Khaki, Samir, et al.
Published: (2025)
by: Khaki, Samir, et al.
Published: (2025)
DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration
by: Zhang, Hanzhi, et al.
Published: (2025)
by: Zhang, Hanzhi, et al.
Published: (2025)
Similar Items
-
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
by: Hooper, Coleman, et al.
Published: (2024) -
SPEED: Speculative Pipelined Execution for Efficient Decoding
by: Hooper, Coleman, et al.
Published: (2023) -
Multipole Attention for Efficient Long Context Reasoning
by: Hooper, Coleman, et al.
Published: (2025) -
ETS: Efficient Tree Search for Inference-Time Scaling
by: Hooper, Coleman, et al.
Published: (2025) -
SqueezeLLM: Dense-and-Sparse Quantization
by: Kim, Sehoon, et al.
Published: (2023)