GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Taneja, Maanas, Shingvi, Purab |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
von: Liu, Guangda, et al.
Veröffentlicht: (2024)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
von: Liu, Zirui, et al.
Veröffentlicht: (2024)
von: Liu, Zirui, et al.
Veröffentlicht: (2024)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)
Improving the Serving Performance of Multi-LoRA Large Language Models via Efficient LoRA and KV Cache Management
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
von: Zhang, Hang, et al.
Veröffentlicht: (2025)
ZO2: Scalable Zeroth-Order Fine-Tuning for Extremely Large Language Models with Limited GPU Memory
von: Wang, Liangyu, et al.
Veröffentlicht: (2025)
von: Wang, Liangyu, et al.
Veröffentlicht: (2025)
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO
von: Barad, Haim, et al.
Veröffentlicht: (2023)
von: Barad, Haim, et al.
Veröffentlicht: (2023)
Accelerating Sparse Ternary GEMM for Quantized ML on Apple Silicon
von: Lipshitz, Baraq, et al.
Veröffentlicht: (2025)
von: Lipshitz, Baraq, et al.
Veröffentlicht: (2025)
LLMPerf: GPU Performance Modeling meets Large Language Models
von: Nguyen, Khoi N. M., et al.
Veröffentlicht: (2025)
von: Nguyen, Khoi N. M., et al.
Veröffentlicht: (2025)
Dual-Signal Adaptive KV-Cache Optimization for Long-Form Video Understanding in Vision-Language Models
von: Sai, Vishnu, et al.
Veröffentlicht: (2026)
von: Sai, Vishnu, et al.
Veröffentlicht: (2026)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
von: Lin, Yujun, et al.
Veröffentlicht: (2024)
von: Lin, Yujun, et al.
Veröffentlicht: (2024)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024)
von: Shkolnik, Moran, et al.
Veröffentlicht: (2024)
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
von: Hendria, Willy Fitra
Veröffentlicht: (2026)
KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoyi, et al.
Veröffentlicht: (2024)
Large-Scale Data Parallelization of Product Quantization and Inverted Indexing Using Dask
von: Abraham, Ashley N., et al.
Veröffentlicht: (2026)
von: Abraham, Ashley N., et al.
Veröffentlicht: (2026)
Model Compression and Efficient Inference for Large Language Models: A Survey
von: Wang, Wenxiao, et al.
Veröffentlicht: (2024)
von: Wang, Wenxiao, et al.
Veröffentlicht: (2024)
Forecasting GPU Performance for Deep Learning Training and Inference
von: Lee, Seonho, et al.
Veröffentlicht: (2024)
von: Lee, Seonho, et al.
Veröffentlicht: (2024)
Efficient GPU implementation of randomized SVD and its applications
von: Struski, Łukasz, et al.
Veröffentlicht: (2021)
von: Struski, Łukasz, et al.
Veröffentlicht: (2021)
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
von: Ma, Xinyue, et al.
Veröffentlicht: (2026)
von: Ma, Xinyue, et al.
Veröffentlicht: (2026)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2024)
von: Mozaffari, Mohammad, et al.
Veröffentlicht: (2024)
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
von: Saha, Rappy, et al.
Veröffentlicht: (2026)
von: Saha, Rappy, et al.
Veröffentlicht: (2026)
SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
EPIC: Efficient Position-Independent Caching for Serving Large Language Models
von: Hu, Junhao, et al.
Veröffentlicht: (2024)
von: Hu, Junhao, et al.
Veröffentlicht: (2024)
PixelBrax: Learning Continuous Control from Pixels End-to-End on the GPU
von: McInroe, Trevor, et al.
Veröffentlicht: (2025)
von: McInroe, Trevor, et al.
Veröffentlicht: (2025)
KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
AutoKernel: Autonomous GPU Kernel Optimization via Iterative Agent-Driven Search
von: Jaber, Jaber, et al.
Veröffentlicht: (2026)
von: Jaber, Jaber, et al.
Veröffentlicht: (2026)
Leveraging Approximate Caching for Faster Retrieval-Augmented Generation
von: Bergman, Shai, et al.
Veröffentlicht: (2025)
von: Bergman, Shai, et al.
Veröffentlicht: (2025)
MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
von: Chitty-Venkata, Krishna Teja, et al.
Veröffentlicht: (2025)
A Performance Evaluation of a Quantized Large Language Model on Various Smartphones
von: Çöplü, Tolga, et al.
Veröffentlicht: (2023)
von: Çöplü, Tolga, et al.
Veröffentlicht: (2023)
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
von: Xue, Leyang, et al.
Veröffentlicht: (2024)
von: Xue, Leyang, et al.
Veröffentlicht: (2024)
Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach
von: Zhang, Yijia, et al.
Veröffentlicht: (2024)
von: Zhang, Yijia, et al.
Veröffentlicht: (2024)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
von: An, Zihao, et al.
Veröffentlicht: (2025)
von: An, Zihao, et al.
Veröffentlicht: (2025)
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
von: Jia, Jinda, et al.
Veröffentlicht: (2026)
von: Jia, Jinda, et al.
Veröffentlicht: (2026)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
von: Li, Zhuojin, et al.
Veröffentlicht: (2025)
VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration
von: Tu, Dezhan, et al.
Veröffentlicht: (2024)
von: Tu, Dezhan, et al.
Veröffentlicht: (2024)
KV-CAR: KV Cache Compression using Autoencoders and KV Reuse in Large Language Models
von: Roy, Sourjya, et al.
Veröffentlicht: (2025)
von: Roy, Sourjya, et al.
Veröffentlicht: (2025)
Fairness in Serving Large Language Models
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
von: Sheng, Ying, et al.
Veröffentlicht: (2023)
RWKV-edge: Deeply Compressed RWKV for Resource-Constrained Devices
von: Choe, Wonkyo, et al.
Veröffentlicht: (2024)
von: Choe, Wonkyo, et al.
Veröffentlicht: (2024)
SENSEi: Input-Sensitive Compilation for Accelerating GNNs
von: Lenadora, Damitha, et al.
Veröffentlicht: (2023)
von: Lenadora, Damitha, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
von: Zhao, Youpeng, et al.
Veröffentlicht: (2024) -
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
von: Liu, Guangda, et al.
Veröffentlicht: (2024) -
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
von: Liu, Zirui, et al.
Veröffentlicht: (2024) -
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
von: Zandieh, Amir, et al.
Veröffentlicht: (2024) -
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
von: Zhou, Zhongzhu, et al.
Veröffentlicht: (2026)