PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Songhao, Lv, Ang, Feng, Xiao, Zhang, Yufei, Zhang, Xun, Yin, Guojun, Lin, Wei, Yan, Rui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PolarQuant: Quantizing KV Caches with Polar Transformation
di: Han, Insu, et al.
Pubblicazione: (2025)
di: Han, Insu, et al.
Pubblicazione: (2025)
PolarQuant: Optimal Gaussian Weight Quantization via Hadamard Rotation for LLM Compression
di: Vicentino, Caio
Pubblicazione: (2026)
di: Vicentino, Caio
Pubblicazione: (2026)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
AffineQuant: Affine Transformation Quantization for Large Language Models
di: Ma, Yuexiao, et al.
Pubblicazione: (2024)
di: Ma, Yuexiao, et al.
Pubblicazione: (2024)
CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
di: Han, Insu, et al.
Pubblicazione: (2025)
di: Han, Insu, et al.
Pubblicazione: (2025)
Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic Tokenization
di: Li, Guanghan, et al.
Pubblicazione: (2024)
di: Li, Guanghan, et al.
Pubblicazione: (2024)
QuantCache: Adaptive Importance-Guided Quantization with Hierarchical Latent and Layer Caching for Video Generation
di: Wu, Junyi, et al.
Pubblicazione: (2025)
di: Wu, Junyi, et al.
Pubblicazione: (2025)
CacheQuant: Comprehensively Accelerated Diffusion Models
di: Liu, Xuewen, et al.
Pubblicazione: (2025)
di: Liu, Xuewen, et al.
Pubblicazione: (2025)
FrameQuant: Flexible Low-Bit Quantization for Transformers
di: Adepu, Harshavardhan, et al.
Pubblicazione: (2024)
di: Adepu, Harshavardhan, et al.
Pubblicazione: (2024)
Beyond Static Testbeds: An Interaction-Centric Agent Simulation Platform for Dynamic Recommender Systems
di: Jin, Song, et al.
Pubblicazione: (2025)
di: Jin, Song, et al.
Pubblicazione: (2025)
Tagging the Thought: Unlocking Personalization Reasoning via Reinforcement Learning
di: Jin, Song, et al.
Pubblicazione: (2025)
di: Jin, Song, et al.
Pubblicazione: (2025)
QuantFace: Efficient Quantization for Face Restoration
di: Li, Jiatong, et al.
Pubblicazione: (2025)
di: Li, Jiatong, et al.
Pubblicazione: (2025)
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
di: Shao, Yuantian, et al.
Pubblicazione: (2025)
di: Shao, Yuantian, et al.
Pubblicazione: (2025)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
di: Chen, Han, et al.
Pubblicazione: (2025)
di: Chen, Han, et al.
Pubblicazione: (2025)
DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
di: Lin, Haokun, et al.
Pubblicazione: (2024)
di: Lin, Haokun, et al.
Pubblicazione: (2024)
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
di: Lu, Haiquan, et al.
Pubblicazione: (2026)
di: Lu, Haiquan, et al.
Pubblicazione: (2026)
Efficient Stochastic Polar Decoder With Correlated Stochastic Computing
di: Li, Jiaxing, et al.
Pubblicazione: (2025)
di: Li, Jiaxing, et al.
Pubblicazione: (2025)
FibQuant: Universal Vector Quantization for Random-Access KV-Cache Compression
di: Lee, Namyoon, et al.
Pubblicazione: (2026)
di: Lee, Namyoon, et al.
Pubblicazione: (2026)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
di: Xiao, Guangxuan, et al.
Pubblicazione: (2022)
di: Xiao, Guangxuan, et al.
Pubblicazione: (2022)
DVD-Quant: Data-free Video Diffusion Transformers Quantization
di: Li, Zhiteng, et al.
Pubblicazione: (2025)
di: Li, Zhiteng, et al.
Pubblicazione: (2025)
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
di: Xi, Haocheng, et al.
Pubblicazione: (2026)
di: Xi, Haocheng, et al.
Pubblicazione: (2026)
High Throughput Polar Code Decoders with Information Bottleneck Quantization
di: Kestel, Claus, et al.
Pubblicazione: (2024)
di: Kestel, Claus, et al.
Pubblicazione: (2024)
Quantized Polarization Redefines Polar Interfaces
di: Pang, Hongsheng, et al.
Pubblicazione: (2025)
di: Pang, Hongsheng, et al.
Pubblicazione: (2025)
ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
di: Zhang, Juntian, et al.
Pubblicazione: (2025)
di: Zhang, Juntian, et al.
Pubblicazione: (2025)
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
di: Liang, Yesheng, et al.
Pubblicazione: (2025)
di: Liang, Yesheng, et al.
Pubblicazione: (2025)
EfficientQuant: An Efficient Post-Training Quantization for CNN-Transformer Hybrid Models on Edge Devices
di: Saha, Shaibal, et al.
Pubblicazione: (2025)
di: Saha, Shaibal, et al.
Pubblicazione: (2025)
QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals
di: Zhang, Nan, et al.
Pubblicazione: (2026)
di: Zhang, Nan, et al.
Pubblicazione: (2026)
QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
di: Feng, Weilun, et al.
Pubblicazione: (2025)
di: Feng, Weilun, et al.
Pubblicazione: (2025)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
di: Duan, Bowen, et al.
Pubblicazione: (2026)
di: Duan, Bowen, et al.
Pubblicazione: (2026)
QuantDemoire: Quantization with Outlier Aware for Image Demoiréing
di: Chen, Zheng, et al.
Pubblicazione: (2025)
di: Chen, Zheng, et al.
Pubblicazione: (2025)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
di: Tao, Wei, et al.
Pubblicazione: (2026)
di: Tao, Wei, et al.
Pubblicazione: (2026)
SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
di: Duanmu, Haojie, et al.
Pubblicazione: (2024)
di: Duanmu, Haojie, et al.
Pubblicazione: (2024)
RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory
di: Zuo, Fei, et al.
Pubblicazione: (2026)
di: Zuo, Fei, et al.
Pubblicazione: (2026)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
di: Lv, Ang, et al.
Pubblicazione: (2024)
di: Lv, Ang, et al.
Pubblicazione: (2024)
Compact Spin-Polarized Positron Acceleration in Multi-Layer Microhole Array Films
di: Dou, Zhen-Ke, et al.
Pubblicazione: (2024)
di: Dou, Zhen-Ke, et al.
Pubblicazione: (2024)
Balanced Low-Complexity and Flexible Error-Correction List Flip Decoding for Polar Codes
di: Lv, Yansong, et al.
Pubblicazione: (2023)
di: Lv, Yansong, et al.
Pubblicazione: (2023)
QuantAttack: Exploiting Dynamic Quantization to Attack Vision Transformers
di: Baras, Amit, et al.
Pubblicazione: (2023)
di: Baras, Amit, et al.
Pubblicazione: (2023)
Low-Complexity PSCL Decoding of Polar Codes
di: Yao, Xinyuanmeng, et al.
Pubblicazione: (2024)
di: Yao, Xinyuanmeng, et al.
Pubblicazione: (2024)
A Universal List Decoding Algorithm with Application to Decoding of Polar Codes
di: Zheng, Xiangping, et al.
Pubblicazione: (2024)
di: Zheng, Xiangping, et al.
Pubblicazione: (2024)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
di: Tang, Hanlin, et al.
Pubblicazione: (2024)
di: Tang, Hanlin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
PolarQuant: Quantizing KV Caches with Polar Transformation
di: Han, Insu, et al.
Pubblicazione: (2025) -
PolarQuant: Optimal Gaussian Weight Quantization via Hadamard Rotation for LLM Compression
di: Vicentino, Caio
Pubblicazione: (2026) -
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025) -
AffineQuant: Affine Transformation Quantization for Large Language Models
di: Ma, Yuexiao, et al.
Pubblicazione: (2024) -
CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
di: Han, Insu, et al.
Pubblicazione: (2025)