PQCache: Product Quantization-based KVCache for Long Context LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Hailin, Ji, Xiaodong, Chen, Yilin, Fu, Fangcheng, Miao, Xupeng, Nie, Xiaonan, Chen, Weipeng, Cui, Bin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025)
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025)
Clover: Regressive Lightweight Speculative Decoding with Sequential Knowledge
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning
von: Wang, Yujie, et al.
Veröffentlicht: (2026)
von: Wang, Yujie, et al.
Veröffentlicht: (2026)
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
von: Ye, Jiancai, et al.
Veröffentlicht: (2026)
von: Ye, Jiancai, et al.
Veröffentlicht: (2026)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
von: Fu, Qichen, et al.
Veröffentlicht: (2024)
AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference
von: He, Zhuomin, et al.
Veröffentlicht: (2025)
von: He, Zhuomin, et al.
Veröffentlicht: (2025)
Clover-2: Accurate Inference for Regressive Lightweight Speculative Decoding
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
DataSculpt: Crafting Data Landscapes for Long-Context LLMs through Multi-Objective Partitioning
von: Lu, Keer, et al.
Veröffentlicht: (2024)
von: Lu, Keer, et al.
Veröffentlicht: (2024)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
von: Lin, Gang, et al.
Veröffentlicht: (2026)
von: Lin, Gang, et al.
Veröffentlicht: (2026)
Accelerating Text-to-Image Editing via Cache-Enabled Sparse Diffusion Inference
von: Yu, Zihao, et al.
Veröffentlicht: (2023)
von: Yu, Zihao, et al.
Veröffentlicht: (2023)
Cocktail: Chunk-Adaptive Mixed-Precision Quantization for Long-Context LLM Inference
von: Tao, Wei, et al.
Veröffentlicht: (2025)
von: Tao, Wei, et al.
Veröffentlicht: (2025)
ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs
von: Ge, Hao, et al.
Veröffentlicht: (2025)
von: Ge, Hao, et al.
Veröffentlicht: (2025)
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
von: Mai, Tho, et al.
Veröffentlicht: (2026)
von: Mai, Tho, et al.
Veröffentlicht: (2026)
FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human Feedback
von: Li, Youquan, et al.
Veröffentlicht: (2024)
von: Li, Youquan, et al.
Veröffentlicht: (2024)
ACON: Optimizing Context Compression for Long-horizon LLM Agents
von: Kang, Minki, et al.
Veröffentlicht: (2025)
von: Kang, Minki, et al.
Veröffentlicht: (2025)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2026)
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
von: Dzikanyanga, Gradwell, et al.
Veröffentlicht: (2026)
von: Dzikanyanga, Gradwell, et al.
Veröffentlicht: (2026)
Improving Automatic Parallel Training via Balanced Memory Workload Optimization
von: Wang, Yujie, et al.
Veröffentlicht: (2023)
von: Wang, Yujie, et al.
Veröffentlicht: (2023)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
von: Yang, Xintong, et al.
Veröffentlicht: (2026)
von: Yang, Xintong, et al.
Veröffentlicht: (2026)
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
von: Lin, Bin, et al.
Veröffentlicht: (2024)
von: Lin, Bin, et al.
Veröffentlicht: (2024)
SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
von: Zhao, Xinye, et al.
Veröffentlicht: (2025)
von: Zhao, Xinye, et al.
Veröffentlicht: (2025)
LLM-based Translation Inference with Iterative Bilingual Understanding
von: Chen, Andong, et al.
Veröffentlicht: (2024)
von: Chen, Andong, et al.
Veröffentlicht: (2024)
Incentivizing In-depth Reasoning over Long Contexts with Process Advantage Shaping
von: Peng, Miao, et al.
Veröffentlicht: (2026)
von: Peng, Miao, et al.
Veröffentlicht: (2026)
KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
von: Wang, Jiahao, et al.
Veröffentlicht: (2025)
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
von: Yu, Hongli, et al.
Veröffentlicht: (2025)
von: Yu, Hongli, et al.
Veröffentlicht: (2025)
Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Tasks
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
TravelBench : Exploring LLM Performance in Low-Resource Domains
von: Billa, Srinivas, et al.
Veröffentlicht: (2025)
von: Billa, Srinivas, et al.
Veröffentlicht: (2025)
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
von: Shi, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Shi, Zhiyuan, et al.
Veröffentlicht: (2026)
Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management
von: Lu, Miao, et al.
Veröffentlicht: (2025)
von: Lu, Miao, et al.
Veröffentlicht: (2025)
Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
von: Shi, Yaorui, et al.
Veröffentlicht: (2025)
von: Shi, Yaorui, et al.
Veröffentlicht: (2025)
Pause-Tuning for Long-Context Comprehension: A Lightweight Approach to LLM Attention Recalibration
von: Begin, James, et al.
Veröffentlicht: (2025)
von: Begin, James, et al.
Veröffentlicht: (2025)
LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
von: Chen, Ziyang, et al.
Veröffentlicht: (2026)
von: Chen, Ziyang, et al.
Veröffentlicht: (2026)
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
Context Memorization for Efficient Long Context Generation
von: Okoshi, Yasuyuki, et al.
Veröffentlicht: (2026)
von: Okoshi, Yasuyuki, et al.
Veröffentlicht: (2026)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
von: Hu, Xing, et al.
Veröffentlicht: (2024)
von: Hu, Xing, et al.
Veröffentlicht: (2024)
LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
von: Jin, Hongye, et al.
Veröffentlicht: (2024)
von: Jin, Hongye, et al.
Veröffentlicht: (2024)
Evaluating Zero-Shot Long-Context LLM Compression
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
von: Wang, Chenyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
von: Ji, Xiaodong, et al.
Veröffentlicht: (2025) -
Clover: Regressive Lightweight Speculative Decoding with Sequential Knowledge
von: Xiao, Bin, et al.
Veröffentlicht: (2024) -
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning
von: Wang, Yujie, et al.
Veröffentlicht: (2026) -
DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference
von: Ye, Jiancai, et al.
Veröffentlicht: (2026) -
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
von: Fu, Qichen, et al.
Veröffentlicht: (2024)