KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zirui, Yuan, Jiayi, Jin, Hongye, Zhong, Shaochen, Xu, Zhaozhuo, Braverman, Vladimir, Chen, Beidi, Hu, Xia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches
von: Yuan, Jiayi, et al.
Veröffentlicht: (2024)
von: Yuan, Jiayi, et al.
Veröffentlicht: (2024)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
von: Liu, Minghui, et al.
Veröffentlicht: (2025)
von: Liu, Minghui, et al.
Veröffentlicht: (2025)
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
von: Bergach, Mohamed Amine
Veröffentlicht: (2026)
von: Bergach, Mohamed Amine
Veröffentlicht: (2026)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
von: Taneja, Maanas, et al.
Veröffentlicht: (2026)
von: Taneja, Maanas, et al.
Veröffentlicht: (2026)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
von: Du, Dayou, et al.
Veröffentlicht: (2025)
von: Du, Dayou, et al.
Veröffentlicht: (2025)
VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration
von: Tu, Dezhan, et al.
Veröffentlicht: (2024)
von: Tu, Dezhan, et al.
Veröffentlicht: (2024)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
von: Lin, Yujun, et al.
Veröffentlicht: (2024)
von: Lin, Yujun, et al.
Veröffentlicht: (2024)
SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving
von: Zhang, Quqing, et al.
Veröffentlicht: (2026)
von: Zhang, Quqing, et al.
Veröffentlicht: (2026)
HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing
von: Liu, Minghui, et al.
Veröffentlicht: (2024)
von: Liu, Minghui, et al.
Veröffentlicht: (2024)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
L1RA: Dynamic Rank Assignment in LoRA Fine-Tuning
von: Singh, Raul, et al.
Veröffentlicht: (2025)
von: Singh, Raul, et al.
Veröffentlicht: (2025)
Unlocking Data-free Low-bit Quantization with Matrix Decomposition for KV Cache Compression
von: Liu, Peiyu, et al.
Veröffentlicht: (2024)
von: Liu, Peiyu, et al.
Veröffentlicht: (2024)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
von: Liu, Hongyao, et al.
Veröffentlicht: (2026)
von: Liu, Hongyao, et al.
Veröffentlicht: (2026)
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
von: Hu, Junhao, et al.
Veröffentlicht: (2024)
von: Hu, Junhao, et al.
Veröffentlicht: (2024)
Systematic Evaluation of Optimization Techniques for Long-Context Language Models
von: Ahmed, Ammar, et al.
Veröffentlicht: (2025)
von: Ahmed, Ammar, et al.
Veröffentlicht: (2025)
Word Salad Chopper: Reasoning Models Waste A Ton Of Decoding Budget On Useless Repetitions, Self-Knowingly
von: Xie, Wenya, et al.
Veröffentlicht: (2025)
von: Xie, Wenya, et al.
Veröffentlicht: (2025)
InnerQ: Hardware-Aware Tuning-Free Quantization of KV Cache for Large Language Models
von: Hosseini, Sayed Mohammadreza Tayaranian, et al.
Veröffentlicht: (2026)
von: Hosseini, Sayed Mohammadreza Tayaranian, et al.
Veröffentlicht: (2026)
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
von: Xu, Pusheng, et al.
Veröffentlicht: (2025)
von: Xu, Pusheng, et al.
Veröffentlicht: (2025)
Input-Gen: Guided Generation of Stateful Inputs for Testing, Tuning, and Training
von: Ivanov, Ivan R., et al.
Veröffentlicht: (2024)
von: Ivanov, Ivan R., et al.
Veröffentlicht: (2024)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
von: Fang, Yunhua, et al.
Veröffentlicht: (2025)
von: Fang, Yunhua, et al.
Veröffentlicht: (2025)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
Accurate KV Cache Quantization with Outlier Tokens Tracing
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO
von: Barad, Haim, et al.
Veröffentlicht: (2023)
von: Barad, Haim, et al.
Veröffentlicht: (2023)
A TRRIP Down Memory Lane: Temperature-Based Re-Reference Interval Prediction For Instruction Caching
von: Kao, Henry, et al.
Veröffentlicht: (2025)
von: Kao, Henry, et al.
Veröffentlicht: (2025)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
von: Tofigh, Mani, et al.
Veröffentlicht: (2025)
von: Tofigh, Mani, et al.
Veröffentlicht: (2025)
Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
von: Wang, Guanchu, et al.
Veröffentlicht: (2024)
von: Wang, Guanchu, et al.
Veröffentlicht: (2024)
Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
von: Dong, Ximing, et al.
Veröffentlicht: (2026)
von: Dong, Ximing, et al.
Veröffentlicht: (2026)
Efficient Hybrid Amplitude-Phase Quantization for Multi-Antenna Relay System
von: Kim, Changdae, et al.
Veröffentlicht: (2025)
von: Kim, Changdae, et al.
Veröffentlicht: (2025)
Stencil-Lifting: Hierarchical Recursive Lifting System for Extracting Summary of Stencil Kernel in Legacy Codes
von: Li, Mingyi, et al.
Veröffentlicht: (2025)
von: Li, Mingyi, et al.
Veröffentlicht: (2025)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
AttentionEngine: A Versatile Framework for Efficient Attention Mechanisms on Diverse Hardware Platforms
von: Chen, Feiyang, et al.
Veröffentlicht: (2025)
von: Chen, Feiyang, et al.
Veröffentlicht: (2025)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
Sirius: Contextual Sparsity with Correction for Efficient LLMs
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2025)
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2025)
SQuat: Subspace-orthogonal KV Cache Quantization
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
von: Jeong, Bodon, et al.
Veröffentlicht: (2026)
von: Jeong, Bodon, et al.
Veröffentlicht: (2026)
KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches
von: Yuan, Jiayi, et al.
Veröffentlicht: (2024) -
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
von: Zandieh, Amir, et al.
Veröffentlicht: (2024) -
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026) -
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
von: Liu, Minghui, et al.
Veröffentlicht: (2025) -
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
von: Bergach, Mohamed Amine
Veröffentlicht: (2026)