DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shao, Yuantian, Chen, Yuanteng, Wang, Peisong, Yu, Jianlin, Lin, Jing, Yao, Yiwu, Wei, Zhihui, Cheng, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Block Rotation is All You Need for MXFP4 Quantization
von: Shao, Yuantian, et al.
Veröffentlicht: (2025)
von: Shao, Yuantian, et al.
Veröffentlicht: (2025)
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
von: Liang, Yesheng, et al.
Veröffentlicht: (2025)
von: Liang, Yesheng, et al.
Veröffentlicht: (2025)
Ban&Pick: Ehancing Performance and Efficiency of MoE-LLMs via Smarter Routing
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
Certain Head, Uncertain Tail: Expert-Sample for Test-Time Scaling in Fine-Grained MoE
von: Chen, Yuanteng, et al.
Veröffentlicht: (2026)
von: Chen, Yuanteng, et al.
Veröffentlicht: (2026)
FlatQuant: Flatness Matters for LLM Quantization
von: Sun, Yuxuan, et al.
Veröffentlicht: (2024)
von: Sun, Yuxuan, et al.
Veröffentlicht: (2024)
DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization
von: Lin, Haokun, et al.
Veröffentlicht: (2026)
von: Lin, Haokun, et al.
Veröffentlicht: (2026)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
PolarQuant: Optimal Gaussian Weight Quantization via Hadamard Rotation for LLM Compression
von: Vicentino, Caio
Veröffentlicht: (2026)
von: Vicentino, Caio
Veröffentlicht: (2026)
DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs
von: Lin, Haokun, et al.
Veröffentlicht: (2024)
von: Lin, Haokun, et al.
Veröffentlicht: (2024)
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration
von: Wu, Songhao, et al.
Veröffentlicht: (2025)
von: Wu, Songhao, et al.
Veröffentlicht: (2025)
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
von: Li, Pingzhi, et al.
Veröffentlicht: (2024)
PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
von: Chen, Mengzhao, et al.
Veröffentlicht: (2024)
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2022)
IntraSlice: Towards High-Performance Structural Pruning with Block-Intra PCA for LLMs
von: Li, Meng, et al.
Veröffentlicht: (2026)
von: Li, Meng, et al.
Veröffentlicht: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
von: Liu, Guangda, et al.
Veröffentlicht: (2025)
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
von: Lu, Haiquan, et al.
Veröffentlicht: (2026)
von: Lu, Haiquan, et al.
Veröffentlicht: (2026)
LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation
von: Chen, Han, et al.
Veröffentlicht: (2025)
von: Chen, Han, et al.
Veröffentlicht: (2025)
VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
von: Wang, Dongwei, et al.
Veröffentlicht: (2026)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
von: Tang, Hanlin, et al.
Veröffentlicht: (2024)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
von: Li, Xing, et al.
Veröffentlicht: (2025)
von: Li, Xing, et al.
Veröffentlicht: (2025)
IsoQuant: Hardware-Aligned SO(4) Isoclinic Rotations for LLM KV Cache Compression
von: Ji, Zhongping
Veröffentlicht: (2026)
von: Ji, Zhongping
Veröffentlicht: (2026)
MobileQuant: Mobile-friendly Quantization for On-device Language Models
von: Tan, Fuwen, et al.
Veröffentlicht: (2024)
von: Tan, Fuwen, et al.
Veröffentlicht: (2024)
FrameQuant: Flexible Low-Bit Quantization for Transformers
von: Adepu, Harshavardhan, et al.
Veröffentlicht: (2024)
von: Adepu, Harshavardhan, et al.
Veröffentlicht: (2024)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
von: Shen, Xuan, et al.
Veröffentlicht: (2023)
von: Shen, Xuan, et al.
Veröffentlicht: (2023)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
GS-Quant: Granular Semantic and Generative Structural Quantization for Knowledge Graph Completion
von: Xie, Qizhuo, et al.
Veröffentlicht: (2026)
von: Xie, Qizhuo, et al.
Veröffentlicht: (2026)
LLMEasyQuant: Scalable Quantization for Parallel and Distributed LLM Inference
von: Liu, Dong, et al.
Veröffentlicht: (2024)
von: Liu, Dong, et al.
Veröffentlicht: (2024)
$C^3$: Confidence Calibration Model Cascade for Inference-Efficient Cross-Lingual Natural Language Understanding
von: Lu, Taixi, et al.
Veröffentlicht: (2024)
von: Lu, Taixi, et al.
Veröffentlicht: (2024)
QUAD: Quantization and Parameter-Efficient Tuning of LLM with Activation Decomposition
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
A Bionic Natural Language Parser Equivalent to a Pushdown Automaton
von: Wei, Zhenghao, et al.
Veröffentlicht: (2024)
von: Wei, Zhenghao, et al.
Veröffentlicht: (2024)
QAQ: Quality Adaptive Quantization for LLM KV Cache
von: Dong, Shichen, et al.
Veröffentlicht: (2024)
von: Dong, Shichen, et al.
Veröffentlicht: (2024)
ReSpinQuant: Efficient Layer-Wise LLM Quantization via Subspace Residual Rotation Approximation
von: Kim, Suyoung, et al.
Veröffentlicht: (2026)
von: Kim, Suyoung, et al.
Veröffentlicht: (2026)
LMStyle Benchmark: Evaluating Text Style Transfer for Chatbots
von: Chen, Jianlin
Veröffentlicht: (2024)
von: Chen, Jianlin
Veröffentlicht: (2024)
Calibrating Reasoning in Language Models with Internal Consistency
von: Xie, Zhihui, et al.
Veröffentlicht: (2024)
von: Xie, Zhihui, et al.
Veröffentlicht: (2024)
Calibrating Beyond English: Language Diversity for Better Quantized Multilingual LLM
von: Chimoto, Everlyn Asiko, et al.
Veröffentlicht: (2026)
von: Chimoto, Everlyn Asiko, et al.
Veröffentlicht: (2026)
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
von: Lin, Ji, et al.
Veröffentlicht: (2023)
von: Lin, Ji, et al.
Veröffentlicht: (2023)
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Block Rotation is All You Need for MXFP4 Quantization
von: Shao, Yuantian, et al.
Veröffentlicht: (2025) -
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025) -
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
von: Liang, Yesheng, et al.
Veröffentlicht: (2025) -
Ban&Pick: Ehancing Performance and Efficiency of MoE-LLMs via Smarter Routing
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025) -
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)