XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
Fuente:
arXiv
Salvato in:
| Autori principali: | Tomar, Aditya, Hooper, Coleman, Lee, Minjae, Xi, Haocheng, Tiwari, Rishabh, Kang, Wonjun, Manolache, Luca, Mahoney, Michael W., Keutzer, Kurt, Gholami, Amir |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
di: Xi, Haocheng, et al.
Pubblicazione: (2026)
di: Xi, Haocheng, et al.
Pubblicazione: (2026)
AI and Memory Wall
di: Gholami, Amir, et al.
Pubblicazione: (2024)
di: Gholami, Amir, et al.
Pubblicazione: (2024)
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
Multipole Attention for Efficient Long Context Reasoning
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
Beyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
di: Kim, Minseo, et al.
Pubblicazione: (2025)
di: Kim, Minseo, et al.
Pubblicazione: (2025)
Reward Under Attack: Analyzing the Robustness and Hackability of Process Reward Models
di: Tiwari, Rishabh, et al.
Pubblicazione: (2026)
di: Tiwari, Rishabh, et al.
Pubblicazione: (2026)
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
di: Maheswaran, Monishwaran, et al.
Pubblicazione: (2025)
di: Maheswaran, Monishwaran, et al.
Pubblicazione: (2025)
SciML Agents: Write the Solver, Not the Solution
di: Gaonkar, Saarth, et al.
Pubblicazione: (2025)
di: Gaonkar, Saarth, et al.
Pubblicazione: (2025)
Residual Context Diffusion Language Models
di: Hu, Yuezhou, et al.
Pubblicazione: (2026)
di: Hu, Yuezhou, et al.
Pubblicazione: (2026)
Squeezed Attention: Accelerating Long Context Length LLM Inference
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
di: Hooper, Coleman, et al.
Pubblicazione: (2024)
ETS: Efficient Tree Search for Inference-Time Scaling
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
SqueezeLLM: Dense-and-Sparse Quantization
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
SPEED: Speculative Pipelined Execution for Efficient Decoding
di: Hooper, Coleman, et al.
Pubblicazione: (2023)
di: Hooper, Coleman, et al.
Pubblicazione: (2023)
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
di: Yang, Haoqi, et al.
Pubblicazione: (2025)
di: Yang, Haoqi, et al.
Pubblicazione: (2025)
CDLM: Consistency Diffusion Language Models For Faster Sampling
di: Kim, Minseo, et al.
Pubblicazione: (2025)
di: Kim, Minseo, et al.
Pubblicazione: (2025)
Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
di: Hooper, Coleman, et al.
Pubblicazione: (2026)
di: Hooper, Coleman, et al.
Pubblicazione: (2026)
Characterizing Prompt Compression Methods for Long Context Inference
di: Jha, Siddharth, et al.
Pubblicazione: (2024)
di: Jha, Siddharth, et al.
Pubblicazione: (2024)
Learned Best-Effort LLM Serving
di: Jha, Siddharth, et al.
Pubblicazione: (2024)
di: Jha, Siddharth, et al.
Pubblicazione: (2024)
Towards Foundation Models for Scientific Machine Learning: Characterizing Scaling and Transfer Behavior
di: Subramanian, Shashank, et al.
Pubblicazione: (2023)
di: Subramanian, Shashank, et al.
Pubblicazione: (2023)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
di: Hooper, Coleman, et al.
Pubblicazione: (2025)
An LLM Compiler for Parallel Function Calling
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
di: Kim, Sehoon, et al.
Pubblicazione: (2023)
TinyAgent: Function Calling at the Edge
di: Erdogan, Lutfi Eren, et al.
Pubblicazione: (2024)
di: Erdogan, Lutfi Eren, et al.
Pubblicazione: (2024)
COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
di: Xi, Haocheng, et al.
Pubblicazione: (2024)
di: Xi, Haocheng, et al.
Pubblicazione: (2024)
Agentic Test-Time Scaling for WebAgents
di: Lee, Nicholas, et al.
Pubblicazione: (2026)
di: Lee, Nicholas, et al.
Pubblicazione: (2026)
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
di: Xi, Haocheng, et al.
Pubblicazione: (2026)
di: Xi, Haocheng, et al.
Pubblicazione: (2026)
CoKV: Optimizing KV Cache Allocation via Cooperative Game
di: Sun, Qiheng, et al.
Pubblicazione: (2025)
di: Sun, Qiheng, et al.
Pubblicazione: (2025)
LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement
di: Lee, Nicholas, et al.
Pubblicazione: (2024)
di: Lee, Nicholas, et al.
Pubblicazione: (2024)
XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference
di: Li, Weizhuo, et al.
Pubblicazione: (2024)
di: Li, Weizhuo, et al.
Pubblicazione: (2024)
The Pitfalls of KV Cache Compression
di: Chen, Alex, et al.
Pubblicazione: (2025)
di: Chen, Alex, et al.
Pubblicazione: (2025)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
di: Yu, Bohan, et al.
Pubblicazione: (2025)
di: Yu, Bohan, et al.
Pubblicazione: (2025)
GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache Eviction
di: Li, Xuelin, et al.
Pubblicazione: (2025)
di: Li, Xuelin, et al.
Pubblicazione: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
di: Liu, Guangda, et al.
Pubblicazione: (2025)
di: Liu, Guangda, et al.
Pubblicazione: (2025)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
di: Ganjihal, Sanjeev Rao
Pubblicazione: (2026)
di: Ganjihal, Sanjeev Rao
Pubblicazione: (2026)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
di: Fang, Yunhua, et al.
Pubblicazione: (2025)
di: Fang, Yunhua, et al.
Pubblicazione: (2025)
Leyline: KV Cache Directives for Agentic Inference
di: Ma, Bole, et al.
Pubblicazione: (2026)
di: Ma, Bole, et al.
Pubblicazione: (2026)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
di: Yao, Jiayi, et al.
Pubblicazione: (2026)
di: Yao, Jiayi, et al.
Pubblicazione: (2026)
AlignedKV: Reducing Memory Access of KV-Cache with Precision-Aligned Quantization
di: Tan, Yifan, et al.
Pubblicazione: (2024)
di: Tan, Yifan, et al.
Pubblicazione: (2024)
NestedKV: Nested Memory Routing for Long-Context KV Cache Compression
di: Chen, Hong, et al.
Pubblicazione: (2026)
di: Chen, Hong, et al.
Pubblicazione: (2026)
KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
di: Nadali, Alireza, et al.
Pubblicazione: (2026)
di: Nadali, Alireza, et al.
Pubblicazione: (2026)
Documenti analoghi
-
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025) -
LoSA: Locality Aware Sparse Attention for Block-Wise Diffusion Language Models
di: Xi, Haocheng, et al.
Pubblicazione: (2026) -
AI and Memory Wall
di: Gholami, Amir, et al.
Pubblicazione: (2024) -
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
di: Hooper, Coleman, et al.
Pubblicazione: (2024) -
Multipole Attention for Efficient Long Context Reasoning
di: Hooper, Coleman, et al.
Pubblicazione: (2025)