QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tiwari, Rishabh, Xi, Haocheng, Tomar, Aditya, Hooper, Coleman, Kim, Sehoon, Horton, Maxwell, Najibi, Mahyar, Mahoney, Michael W., Keutzer, Kurt, Gholami, Amir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
von: Tomar, Aditya, et al.
Veröffentlicht: (2025)
von: Tomar, Aditya, et al.
Veröffentlicht: (2025)
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
von: Hooper, Coleman, et al.
Veröffentlicht: (2024)
von: Hooper, Coleman, et al.
Veröffentlicht: (2024)
SPEED: Speculative Pipelined Execution for Efficient Decoding
von: Hooper, Coleman, et al.
Veröffentlicht: (2023)
von: Hooper, Coleman, et al.
Veröffentlicht: (2023)
AI and Memory Wall
von: Gholami, Amir, et al.
Veröffentlicht: (2024)
von: Gholami, Amir, et al.
Veröffentlicht: (2024)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
von: Maheswaran, Monishwaran, et al.
Veröffentlicht: (2025)
von: Maheswaran, Monishwaran, et al.
Veröffentlicht: (2025)
SqueezeLLM: Dense-and-Sparse Quantization
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2026)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2026)
Multipole Attention for Efficient Long Context Reasoning
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
Squeezed Attention: Accelerating Long Context Length LLM Inference
von: Hooper, Coleman, et al.
Veröffentlicht: (2024)
von: Hooper, Coleman, et al.
Veröffentlicht: (2024)
Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
von: Hooper, Coleman, et al.
Veröffentlicht: (2026)
von: Hooper, Coleman, et al.
Veröffentlicht: (2026)
SciML Agents: Write the Solver, Not the Solution
von: Gaonkar, Saarth, et al.
Veröffentlicht: (2025)
von: Gaonkar, Saarth, et al.
Veröffentlicht: (2025)
Learned Best-Effort LLM Serving
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)
PolarQuant: Quantizing KV Caches with Polar Transformation
von: Han, Insu, et al.
Veröffentlicht: (2025)
von: Han, Insu, et al.
Veröffentlicht: (2025)
ETS: Efficient Tree Search for Inference-Time Scaling
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
Residual Context Diffusion Language Models
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2026)
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
An LLM Compiler for Parallel Function Calling
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
von: Kim, Sehoon, et al.
Veröffentlicht: (2023)
Characterizing Prompt Compression Methods for Long Context Inference
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)
von: Jha, Siddharth, et al.
Veröffentlicht: (2024)
TinyAgent: Function Calling at the Edge
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2024)
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2024)
CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
von: Han, Insu, et al.
Veröffentlicht: (2025)
von: Han, Insu, et al.
Veröffentlicht: (2025)
CDLM: Consistency Diffusion Language Models For Faster Sampling
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
von: Kim, Minseo, et al.
Veröffentlicht: (2025)
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
von: Shukla, Shikhar
Veröffentlicht: (2026)
von: Shukla, Shikhar
Veröffentlicht: (2026)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
von: Lee, Namyoon, et al.
Veröffentlicht: (2026)
von: Lee, Namyoon, et al.
Veröffentlicht: (2026)
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement
von: Lee, Nicholas, et al.
Veröffentlicht: (2024)
von: Lee, Nicholas, et al.
Veröffentlicht: (2024)
Efficient and Scalable Estimation of Tool Representations in Vector Space
von: Moon, Suhong, et al.
Veröffentlicht: (2024)
von: Moon, Suhong, et al.
Veröffentlicht: (2024)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
von: Hooper, Coleman, et al.
Veröffentlicht: (2025)
ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
von: Chen, Han, et al.
Veröffentlicht: (2025)
von: Chen, Han, et al.
Veröffentlicht: (2025)
RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory
von: Zuo, Fei, et al.
Veröffentlicht: (2026)
von: Zuo, Fei, et al.
Veröffentlicht: (2026)
Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior
von: Subramanian, Shashank, et al.
Veröffentlicht: (2023)
von: Subramanian, Shashank, et al.
Veröffentlicht: (2023)
Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2025)
von: Erdogan, Lutfi Eren, et al.
Veröffentlicht: (2025)
$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration
von: Wu, Songhao, et al.
Veröffentlicht: (2025)
von: Wu, Songhao, et al.
Veröffentlicht: (2025)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
von: Tao, Wei, et al.
Veröffentlicht: (2026)
von: Tao, Wei, et al.
Veröffentlicht: (2026)
SpecMemo: Speculative Decoding is in Your Pocket
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
von: Tomar, Aditya, et al.
Veröffentlicht: (2025) -
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
von: Hooper, Coleman, et al.
Veröffentlicht: (2024) -
SPEED: Speculative Pipelined Execution for Efficient Decoding
von: Hooper, Coleman, et al.
Veröffentlicht: (2023) -
AI and Memory Wall
von: Gholami, Amir, et al.
Veröffentlicht: (2024) -
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
von: Maheswaran, Monishwaran, et al.
Veröffentlicht: (2025)