Low-Rank Correction for Quantized LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Scetbon, Meyer, Hensman, James |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension
di: Gong, Wenbo, et al.
Pubblicazione: (2025)
di: Gong, Wenbo, et al.
Pubblicazione: (2025)
Pyramid Vector Quantization for LLMs
di: van der Ouderaa, Tycho F. A., et al.
Pubblicazione: (2024)
di: van der Ouderaa, Tycho F. A., et al.
Pubblicazione: (2024)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
di: Gadhikar, Advait, et al.
Pubblicazione: (2025)
di: Gadhikar, Advait, et al.
Pubblicazione: (2025)
LQER: Low-Rank Quantization Error Reconstruction for LLMs
di: Zhang, Cheng, et al.
Pubblicazione: (2024)
di: Zhang, Cheng, et al.
Pubblicazione: (2024)
Low-Rank Quantization-Aware Training for LLMs
di: Bondarenko, Yelysei, et al.
Pubblicazione: (2024)
di: Bondarenko, Yelysei, et al.
Pubblicazione: (2024)
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
di: Ma, Chao, et al.
Pubblicazione: (2024)
di: Ma, Chao, et al.
Pubblicazione: (2024)
Gradient Multi-Normalization for Stateless and Scalable LLM Training
di: Scetbon, Meyer, et al.
Pubblicazione: (2025)
di: Scetbon, Meyer, et al.
Pubblicazione: (2025)
A Fixed-Point Approach for Causal Generative Modeling
di: Scetbon, Meyer, et al.
Pubblicazione: (2024)
di: Scetbon, Meyer, et al.
Pubblicazione: (2024)
Amortized Inference of Causal Models via Conditional Fixed-Point Iterations
di: Mahajan, Divyat, et al.
Pubblicazione: (2024)
di: Mahajan, Divyat, et al.
Pubblicazione: (2024)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
di: Cho, Yoonjun, et al.
Pubblicazione: (2026)
di: Cho, Yoonjun, et al.
Pubblicazione: (2026)
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
di: Kang, Hao, et al.
Pubblicazione: (2024)
di: Kang, Hao, et al.
Pubblicazione: (2024)
LoQT: Low-Rank Adapters for Quantized Pretraining
di: Loeschcke, Sebastian, et al.
Pubblicazione: (2024)
di: Loeschcke, Sebastian, et al.
Pubblicazione: (2024)
DiSK: A Diffusion Model for Structured Knowledge
di: Kitouni, Ouail, et al.
Pubblicazione: (2023)
di: Kitouni, Ouail, et al.
Pubblicazione: (2023)
FLRQ: Faster LLM Quantization with Flexible Low-Rank Matrix Sketching
di: Gul, Hongyaoxing, et al.
Pubblicazione: (2026)
di: Gul, Hongyaoxing, et al.
Pubblicazione: (2026)
SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization
di: Park, Yeonsik, et al.
Pubblicazione: (2026)
di: Park, Yeonsik, et al.
Pubblicazione: (2026)
LoRaQ: Optimized Low Rank Approximation for 4-bit Quantization
di: Bouquet, Yann, et al.
Pubblicazione: (2026)
di: Bouquet, Yann, et al.
Pubblicazione: (2026)
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
di: Ashkboos, Saleh, et al.
Pubblicazione: (2024)
di: Ashkboos, Saleh, et al.
Pubblicazione: (2024)
GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs
di: An, Selim, et al.
Pubblicazione: (2026)
di: An, Selim, et al.
Pubblicazione: (2026)
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
di: Xu, Zifei, et al.
Pubblicazione: (2024)
di: Xu, Zifei, et al.
Pubblicazione: (2024)
CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs
di: Zhou, Zhaojing, et al.
Pubblicazione: (2025)
di: Zhou, Zhaojing, et al.
Pubblicazione: (2025)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
LoPRo: Enhancing Low-Rank Quantization via Permuted Block-Wise Rotation
di: Gu, Hongyaoxing, et al.
Pubblicazione: (2026)
di: Gu, Hongyaoxing, et al.
Pubblicazione: (2026)
TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D Tiling
di: Gu, Hongyaoxing, et al.
Pubblicazione: (2026)
di: Gu, Hongyaoxing, et al.
Pubblicazione: (2026)
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
di: Huang, Beichen, et al.
Pubblicazione: (2025)
di: Huang, Beichen, et al.
Pubblicazione: (2025)
LCQ: Low-Rank Codebook based Quantization for Large Language Models
di: Cai, Wen-Pu, et al.
Pubblicazione: (2024)
di: Cai, Wen-Pu, et al.
Pubblicazione: (2024)
CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
Training Acceleration of Low-Rank Decomposed Networks using Sequential Freezing and Rank Quantization
di: Hajimolahoseini, Habib, et al.
Pubblicazione: (2023)
di: Hajimolahoseini, Habib, et al.
Pubblicazione: (2023)
Revisiting Transformer Layer Parameterization Through Causal Energy Minimization
di: Xu, Jin, et al.
Pubblicazione: (2026)
di: Xu, Jin, et al.
Pubblicazione: (2026)
MLoRQ: Bridging Low-Rank and Quantization for Transformer Compression
di: Gordon, Ofir, et al.
Pubblicazione: (2025)
di: Gordon, Ofir, et al.
Pubblicazione: (2025)
Can Low-Rank Knowledge Distillation in LLMs be Useful for Microelectronic Reasoning?
di: Rouf, Nirjhor, et al.
Pubblicazione: (2024)
di: Rouf, Nirjhor, et al.
Pubblicazione: (2024)
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
di: Xu, Yinggan, et al.
Pubblicazione: (2026)
di: Xu, Yinggan, et al.
Pubblicazione: (2026)
Low-Rank Key Value Attention
di: O'Neill, James, et al.
Pubblicazione: (2026)
di: O'Neill, James, et al.
Pubblicazione: (2026)
FinLoRA: Finetuning Quantized Financial Large Language Models Using Low-Rank Adaptation
di: Wang, Dannong, et al.
Pubblicazione: (2024)
di: Wang, Dannong, et al.
Pubblicazione: (2024)
Efficient Pareto Manifold Learning with Low-Rank Structure
di: Chen, Weiyu, et al.
Pubblicazione: (2024)
di: Chen, Weiyu, et al.
Pubblicazione: (2024)
Scalable Model-Based Clustering with Sequential Monte Carlo
di: Trojan, Connie, et al.
Pubblicazione: (2026)
di: Trojan, Connie, et al.
Pubblicazione: (2026)
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
di: Saxena, Utkarsh, et al.
Pubblicazione: (2024)
di: Saxena, Utkarsh, et al.
Pubblicazione: (2024)
Breaking the Blocks: Continuous Low-Rank Decomposed Scaling for Unified LLM Quantization and Adaptation
di: Tang, Pingzhi, et al.
Pubblicazione: (2026)
di: Tang, Pingzhi, et al.
Pubblicazione: (2026)
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
di: Zhang, Zhenyu, et al.
Pubblicazione: (2024)
di: Zhang, Zhenyu, et al.
Pubblicazione: (2024)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
di: Lee, Seoungsub, et al.
Pubblicazione: (2026)
di: Lee, Seoungsub, et al.
Pubblicazione: (2026)
LoRAP: Low-Rank Aggregation Prompting for Quantized Graph Neural Networks Training
di: Liu, Chenyu, et al.
Pubblicazione: (2026)
di: Liu, Chenyu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension
di: Gong, Wenbo, et al.
Pubblicazione: (2025) -
Pyramid Vector Quantization for LLMs
di: van der Ouderaa, Tycho F. A., et al.
Pubblicazione: (2024) -
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
di: Gadhikar, Advait, et al.
Pubblicazione: (2025) -
LQER: Low-Rank Quantization Error Reconstruction for LLMs
di: Zhang, Cheng, et al.
Pubblicazione: (2024) -
Low-Rank Quantization-Aware Training for LLMs
di: Bondarenko, Yelysei, et al.
Pubblicazione: (2024)