Guardado en:
| Autores principales: | Dong, Shichen, Cheng, Wen, Qin, Jiayu, Wang, Wei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2403.04643 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences
por: Qin, Ziran, et al.
Publicado: (2025)
por: Qin, Ziran, et al.
Publicado: (2025)
KV-CoRE: Benchmarking Data-Dependent Low-Rank Compressibility of KV-Caches in LLMs
por: Chen, Jian, et al.
Publicado: (2026)
por: Chen, Jian, et al.
Publicado: (2026)
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
por: Zuo, Youhui, et al.
Publicado: (2025)
por: Zuo, Youhui, et al.
Publicado: (2025)
AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models
por: Su, Zunhai, et al.
Publicado: (2025)
por: Su, Zunhai, et al.
Publicado: (2025)
Accurate KV Cache Quantization with Outlier Tokens Tracing
por: Su, Yi, et al.
Publicado: (2025)
por: Su, Yi, et al.
Publicado: (2025)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
por: Su, Zunhai, et al.
Publicado: (2025)
por: Su, Zunhai, et al.
Publicado: (2025)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
por: Sun, Hanshi, et al.
Publicado: (2024)
por: Sun, Hanshi, et al.
Publicado: (2024)
VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
por: Yao, Dingyu, et al.
Publicado: (2025)
por: Yao, Dingyu, et al.
Publicado: (2025)
QAQ: Bidirectional Semantic Coherence for Selecting High-Quality Synthetic Code Instructions
por: Lei, Jiayin, et al.
Publicado: (2026)
por: Lei, Jiayin, et al.
Publicado: (2026)
SQuat: Subspace-orthogonal KV Cache Quantization
por: Wang, Hao, et al.
Publicado: (2025)
por: Wang, Hao, et al.
Publicado: (2025)
Unlocking Data-free Low-bit Quantization with Matrix Decomposition for KV Cache Compression
por: Liu, Peiyu, et al.
Publicado: (2024)
por: Liu, Peiyu, et al.
Publicado: (2024)
R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
por: Cai, Zefan, et al.
Publicado: (2025)
por: Cai, Zefan, et al.
Publicado: (2025)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
por: Feng, Yuan, et al.
Publicado: (2024)
por: Feng, Yuan, et al.
Publicado: (2024)
CommVQ: Commutative Vector Quantization for KV Cache Compression
por: Li, Junyan, et al.
Publicado: (2025)
por: Li, Junyan, et al.
Publicado: (2025)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
por: Liu, Xiang, et al.
Publicado: (2025)
por: Liu, Xiang, et al.
Publicado: (2025)
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs
por: Zhou, Xiabin, et al.
Publicado: (2024)
por: Zhou, Xiabin, et al.
Publicado: (2024)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
por: Yu, Bohan, et al.
Publicado: (2025)
por: Yu, Bohan, et al.
Publicado: (2025)
Quantization Dominates Rank Reduction for KV-Cache Compression
por: Salfati, Samuel
Publicado: (2026)
por: Salfati, Samuel
Publicado: (2026)
One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
por: Lu, Liming, et al.
Publicado: (2026)
por: Lu, Liming, et al.
Publicado: (2026)
Efficient Long-Context LLM Inference via KV Cache Clustering
por: Hu, Jie, et al.
Publicado: (2025)
por: Hu, Jie, et al.
Publicado: (2025)
NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics
por: Cai, Zhihang, et al.
Publicado: (2025)
por: Cai, Zhihang, et al.
Publicado: (2025)
Taming the Fragility of KV Cache Eviction in LLM Inference
por: Feng, Yuan, et al.
Publicado: (2025)
por: Feng, Yuan, et al.
Publicado: (2025)
KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs
por: Su, Zunhai, et al.
Publicado: (2025)
por: Su, Zunhai, et al.
Publicado: (2025)
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
por: Liu, Tengxuan, et al.
Publicado: (2025)
por: Liu, Tengxuan, et al.
Publicado: (2025)
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
por: Yang, Dongjie, et al.
Publicado: (2024)
por: Yang, Dongjie, et al.
Publicado: (2024)
dKV-Cache: The Cache for Diffusion Language Models
por: Ma, Xinyin, et al.
Publicado: (2025)
por: Ma, Xinyin, et al.
Publicado: (2025)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
por: Liao, Mengqi, et al.
Publicado: (2025)
por: Liao, Mengqi, et al.
Publicado: (2025)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
por: Dzikanyanga, Gradwell, et al.
Publicado: (2026)
por: Dzikanyanga, Gradwell, et al.
Publicado: (2026)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
por: Cai, Zefan, et al.
Publicado: (2024)
por: Cai, Zefan, et al.
Publicado: (2024)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
por: Liu, Guangda, et al.
Publicado: (2025)
por: Liu, Guangda, et al.
Publicado: (2025)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
por: Tao, Wei, et al.
Publicado: (2026)
por: Tao, Wei, et al.
Publicado: (2026)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
por: Tian, Yuxuan, et al.
Publicado: (2025)
por: Tian, Yuxuan, et al.
Publicado: (2025)
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
por: Guo, Jinyu, et al.
Publicado: (2026)
por: Guo, Jinyu, et al.
Publicado: (2026)
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
por: Yang, Haoqi, et al.
Publicado: (2025)
por: Yang, Haoqi, et al.
Publicado: (2025)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
por: Zandieh, Amir, et al.
Publicado: (2024)
por: Zandieh, Amir, et al.
Publicado: (2024)
HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse
por: An, Yuwei, et al.
Publicado: (2025)
por: An, Yuwei, et al.
Publicado: (2025)
SABlock: Semantic-Aware KV Cache Eviction with Adaptive Compression Block Size
por: Chen, Jinhan, et al.
Publicado: (2025)
por: Chen, Jinhan, et al.
Publicado: (2025)
DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
por: Dehghanighobadi, Zahra, et al.
Publicado: (2026)
por: Dehghanighobadi, Zahra, et al.
Publicado: (2026)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
por: Liu, Zirui, et al.
Publicado: (2024)
por: Liu, Zirui, et al.
Publicado: (2024)
KVzap: Fast, Adaptive, and Faithful KV Cache Pruning
por: Jegou, Simon, et al.
Publicado: (2026)
por: Jegou, Simon, et al.
Publicado: (2026)
Ejemplares similares
-
CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences
por: Qin, Ziran, et al.
Publicado: (2025) -
KV-CoRE: Benchmarking Data-Dependent Low-Rank Compressibility of KV-Caches in LLMs
por: Chen, Jian, et al.
Publicado: (2026) -
WindowKV: Task-Adaptive Group-Wise KV Cache Window Selection for Efficient LLM Inference
por: Zuo, Youhui, et al.
Publicado: (2025) -
AKVQ-VL: Attention-Aware KV Cache Adaptive 2-Bit Quantization for Vision-Language Models
por: Su, Zunhai, et al.
Publicado: (2025) -
Accurate KV Cache Quantization with Outlier Tokens Tracing
por: Su, Yi, et al.
Publicado: (2025)