QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
Fuente:
arXiv
Guardado en:
| Autores principales: | Zandieh, Amir, Daliri, Majid, Han, Insu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PolarQuant: Quantizing KV Caches with Polar Transformation
por: Han, Insu, et al.
Publicado: (2025)
por: Han, Insu, et al.
Publicado: (2025)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
por: Lin, Yujun, et al.
Publicado: (2024)
por: Lin, Yujun, et al.
Publicado: (2024)
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
por: Zandieh, Amir, et al.
Publicado: (2025)
por: Zandieh, Amir, et al.
Publicado: (2025)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
por: Liu, Zirui, et al.
Publicado: (2024)
por: Liu, Zirui, et al.
Publicado: (2024)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
por: Chen, Han, et al.
Publicado: (2025)
por: Chen, Han, et al.
Publicado: (2025)
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
por: Zhou, Zhongzhu, et al.
Publicado: (2026)
por: Zhou, Zhongzhu, et al.
Publicado: (2026)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
por: Su, Zunhai, et al.
Publicado: (2025)
por: Su, Zunhai, et al.
Publicado: (2025)
SQuat: Subspace-orthogonal KV Cache Quantization
por: Wang, Hao, et al.
Publicado: (2025)
por: Wang, Hao, et al.
Publicado: (2025)
ECO: Quantized Training without Full-Precision Master Weights
por: Nikdan, Mahdi, et al.
Publicado: (2026)
por: Nikdan, Mahdi, et al.
Publicado: (2026)
Unlocking the Theory Behind Scaling 1-Bit Neural Networks
por: Daliri, Majid, et al.
Publicado: (2024)
por: Daliri, Majid, et al.
Publicado: (2024)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
por: Chiang, Hung-Yueh, et al.
Publicado: (2025)
por: Chiang, Hung-Yueh, et al.
Publicado: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
por: Du, Dayou, et al.
Publicado: (2025)
por: Du, Dayou, et al.
Publicado: (2025)
Quantization Dominates Rank Reduction for KV-Cache Compression
por: Salfati, Samuel
Publicado: (2026)
por: Salfati, Samuel
Publicado: (2026)
HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing
por: Liu, Minghui, et al.
Publicado: (2024)
por: Liu, Minghui, et al.
Publicado: (2024)
ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
por: Liu, Guangda, et al.
Publicado: (2024)
por: Liu, Guangda, et al.
Publicado: (2024)
AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations
por: Tao, Qian, et al.
Publicado: (2024)
por: Tao, Qian, et al.
Publicado: (2024)
NQKV: A KV Cache Quantization Scheme Based on Normal Distribution Characteristics
por: Cai, Zhihang, et al.
Publicado: (2025)
por: Cai, Zhihang, et al.
Publicado: (2025)
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
por: Aggarwal, Shivam, et al.
Publicado: (2023)
por: Aggarwal, Shivam, et al.
Publicado: (2023)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
por: Liu, Minghui, et al.
Publicado: (2025)
por: Liu, Minghui, et al.
Publicado: (2025)
SLiM: One-shot Quantization and Sparsity with Low-rank Approximation for LLM Weight Compression
por: Mozaffari, Mohammad, et al.
Publicado: (2024)
por: Mozaffari, Mohammad, et al.
Publicado: (2024)
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching
por: Zhao, Youpeng, et al.
Publicado: (2024)
por: Zhao, Youpeng, et al.
Publicado: (2024)
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO
por: Barad, Haim, et al.
Publicado: (2023)
por: Barad, Haim, et al.
Publicado: (2023)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
por: Shkolnik, Moran, et al.
Publicado: (2024)
por: Shkolnik, Moran, et al.
Publicado: (2024)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
por: Lee, Banseok, et al.
Publicado: (2025)
por: Lee, Banseok, et al.
Publicado: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
por: Li, Xing, et al.
Publicado: (2025)
por: Li, Xing, et al.
Publicado: (2025)
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon
por: Bergach, Mohamed Amine
Publicado: (2026)
por: Bergach, Mohamed Amine
Publicado: (2026)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
por: Chu, Kexin, et al.
Publicado: (2025)
por: Chu, Kexin, et al.
Publicado: (2025)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
por: Taneja, Maanas, et al.
Publicado: (2026)
por: Taneja, Maanas, et al.
Publicado: (2026)
RDKV: Rate-Distortion Bit Allocation for Joint Eviction and Quantization of the KV Cache
por: Zhang, Junkai, et al.
Publicado: (2026)
por: Zhang, Junkai, et al.
Publicado: (2026)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
por: Chhugani, Jatin, et al.
Publicado: (2026)
por: Chhugani, Jatin, et al.
Publicado: (2026)
Dual-Signal Adaptive KV-Cache Optimization for Long-Form Video Understanding in Vision-Language Models
por: Sai, Vishnu, et al.
Publicado: (2026)
por: Sai, Vishnu, et al.
Publicado: (2026)
Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents
por: Zhang, Qizheng, et al.
Publicado: (2025)
por: Zhang, Qizheng, et al.
Publicado: (2025)
Zero-Shot, But at What Cost? Unveiling the Hidden Overhead of MILS's LLM-CLIP Framework for Image Captioning
por: Benhammou, Yassir, et al.
Publicado: (2025)
por: Benhammou, Yassir, et al.
Publicado: (2025)
SparKV: Overhead-Aware KV Cache Loading for Efficient On-Device LLM Inference
por: Liu, Hongyao, et al.
Publicado: (2026)
por: Liu, Hongyao, et al.
Publicado: (2026)
KV Cache Transform Coding for Compact Storage in LLM Inference
por: Staniszewski, Konrad, et al.
Publicado: (2025)
por: Staniszewski, Konrad, et al.
Publicado: (2025)
Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
por: Du, Dayou, et al.
Publicado: (2024)
por: Du, Dayou, et al.
Publicado: (2024)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
por: Wang, Dongwei, et al.
Publicado: (2026)
por: Wang, Dongwei, et al.
Publicado: (2026)
Evaluating the Efficacy of Foundational Models: Advancing Benchmarking Practices to Enhance Fine-Tuning Decision-Making
por: Amujo, Oluyemi Enoch, et al.
Publicado: (2024)
por: Amujo, Oluyemi Enoch, et al.
Publicado: (2024)
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
por: Zhao, Qihao, et al.
Publicado: (2024)
por: Zhao, Qihao, et al.
Publicado: (2024)
Model Compression and Efficient Inference for Large Language Models: A Survey
por: Wang, Wenxiao, et al.
Publicado: (2024)
por: Wang, Wenxiao, et al.
Publicado: (2024)
Ejemplares similares
-
PolarQuant: Quantizing KV Caches with Polar Transformation
por: Han, Insu, et al.
Publicado: (2025) -
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
por: Lin, Yujun, et al.
Publicado: (2024) -
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
por: Zandieh, Amir, et al.
Publicado: (2025) -
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
por: Liu, Zirui, et al.
Publicado: (2024) -
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
por: Chen, Han, et al.
Publicado: (2025)