Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Cho, Yoonjun, Jeon, Dongjae, Kim, Soeun, Jeon, Moongyu, No, Albert |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
by: Cho, Yoonjun, et al.
Published: (2025)
by: Cho, Yoonjun, et al.
Published: (2025)
DAPD: Dependency-Aware Parallel Decoding via Attention for Diffusion LLMs
by: Kim, Bumjun, et al.
Published: (2026)
by: Kim, Bumjun, et al.
Published: (2026)
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
by: Yoon, Sangyeon, et al.
Published: (2026)
by: Yoon, Sangyeon, et al.
Published: (2026)
Understanding and Mitigating Memorization in Generative Models via Sharpness of Probability Landscapes
by: Jeon, Dongjae, et al.
Published: (2024)
by: Jeon, Dongjae, et al.
Published: (2024)
Information-Theoretic Discrete Diffusion
by: Jeon, Moongyu, et al.
Published: (2025)
by: Jeon, Moongyu, et al.
Published: (2025)
Two-Stage Grid Optimization for Group-wise Quantization of LLMs
by: Kim, Junhan, et al.
Published: (2026)
by: Kim, Junhan, et al.
Published: (2026)
An Information Theoretic Evaluation Metric For Strong Unlearning
by: Jeon, Dongjae, et al.
Published: (2024)
by: Jeon, Dongjae, et al.
Published: (2024)
BoA: Attention-aware Post-training Quantization without Backpropagation
by: Kim, Junhan, et al.
Published: (2024)
by: Kim, Junhan, et al.
Published: (2024)
Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
by: Kim, Junhan, et al.
Published: (2024)
by: Kim, Junhan, et al.
Published: (2024)
Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR
by: Kim, Soeun, et al.
Published: (2026)
by: Kim, Soeun, et al.
Published: (2026)
GlowQ: Group-Shared LOw-Rank Approximation for Quantized LLMs
by: An, Selim, et al.
Published: (2026)
by: An, Selim, et al.
Published: (2026)
A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
by: Jeung, Wonje, et al.
Published: (2025)
by: Jeung, Wonje, et al.
Published: (2025)
TTA-DAME: Test-Time Adaptation with Domain Augmentation and Model Ensemble for Dynamic Driving Conditions
by: Jeon, Dongjae, et al.
Published: (2025)
by: Jeon, Dongjae, et al.
Published: (2025)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025)
by: Kim, Jiyoon, et al.
Published: (2025)
Low-Rank Quantization-Aware Training for LLMs
by: Bondarenko, Yelysei, et al.
Published: (2024)
by: Bondarenko, Yelysei, et al.
Published: (2024)
ASER: Activation Smoothing and Error Reconstruction for Large Language Model Quantization
by: Zhao, Weibo, et al.
Published: (2024)
by: Zhao, Weibo, et al.
Published: (2024)
FAQ: Mitigating Quantization Error via Regenerating Calibration Data with Family-Aware Quantization
by: Xiao, Haiyang, et al.
Published: (2026)
by: Xiao, Haiyang, et al.
Published: (2026)
Rainbow Padding: Mitigating Early Termination in Instruction-Tuned Diffusion LLMs
by: Kim, Bumjun, et al.
Published: (2025)
by: Kim, Bumjun, et al.
Published: (2025)
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy
by: Lee, Geonho, et al.
Published: (2024)
by: Lee, Geonho, et al.
Published: (2024)
Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
by: Chhugani, Jatin, et al.
Published: (2026)
by: Chhugani, Jatin, et al.
Published: (2026)
MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition
by: Lee, Seoungsub, et al.
Published: (2026)
by: Lee, Seoungsub, et al.
Published: (2026)
Edge-FIT: Federated Instruction Tuning of Quantized LLMs for Privacy-Preserving Smart Home Environments
by: Venkatesh, Vinay, et al.
Published: (2025)
by: Venkatesh, Vinay, et al.
Published: (2025)
Dissecting Quantization Error: A Concentration-Alignment Perspective
by: Federici, Marco, et al.
Published: (2026)
by: Federici, Marco, et al.
Published: (2026)
Enhancing Delta Compression in LLMs via SVD-based Quantization Error Minimization
by: Xiong, Boya, et al.
Published: (2025)
by: Xiong, Boya, et al.
Published: (2025)
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
by: Xu, Yinggan, et al.
Published: (2026)
by: Xu, Yinggan, et al.
Published: (2026)
On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs
by: Ye, Rongguang, et al.
Published: (2025)
by: Ye, Rongguang, et al.
Published: (2025)
Multi-Level Knowledge Distillation and Dynamic Self-Supervised Learning for Continual Learning
by: Kim, Taeheon, et al.
Published: (2025)
by: Kim, Taeheon, et al.
Published: (2025)
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
LQER: Low-Rank Quantization Error Reconstruction for LLMs
by: Zhang, Cheng, et al.
Published: (2024)
by: Zhang, Cheng, et al.
Published: (2024)
Interpreting the Effects of Quantization on LLMs
by: Singh, Manpreet, et al.
Published: (2025)
by: Singh, Manpreet, et al.
Published: (2025)
Runtime-Certified Bounded-Error Quantized Attention
by: Calver, Dean
Published: (2026)
by: Calver, Dean
Published: (2026)
QUAIL: Quantization Aware Unlearning for Mitigating Misinformation in LLMs
by: Mishra, Himanshu, et al.
Published: (2026)
by: Mishra, Himanshu, et al.
Published: (2026)
A Theoretical Analysis of Why Masked Diffusion Models Mitigate the Reversal Curse
by: Jeon, Moongyu, et al.
Published: (2026)
by: Jeon, Moongyu, et al.
Published: (2026)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
by: Shkolnik, Moran, et al.
Published: (2024)
by: Shkolnik, Moran, et al.
Published: (2024)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
by: Park, Gunho, et al.
Published: (2025)
by: Park, Gunho, et al.
Published: (2025)
Quantization Dominates Rank Reduction for KV-Cache Compression
by: Salfati, Samuel
Published: (2026)
by: Salfati, Samuel
Published: (2026)
APreQEL: Adaptive Mixed Precision Quantization For Edge LLMs
by: Bouzouad, Meriem, et al.
Published: (2026)
by: Bouzouad, Meriem, et al.
Published: (2026)
EasyQuant: An Efficient Data-free Quantization Algorithm for LLMs
by: Tang, Hanlin, et al.
Published: (2024)
by: Tang, Hanlin, et al.
Published: (2024)
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025)
by: Chmiel, Brian, et al.
Published: (2025)
Early Quantization Shrinks Codebook: A Simple Fix for Diversity-Preserving Tokenization
by: Zhao, Wenhao, et al.
Published: (2026)
by: Zhao, Wenhao, et al.
Published: (2026)
Similar Items
-
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
by: Cho, Yoonjun, et al.
Published: (2025) -
DAPD: Dependency-Aware Parallel Decoding via Attention for Diffusion LLMs
by: Kim, Bumjun, et al.
Published: (2026) -
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
by: Yoon, Sangyeon, et al.
Published: (2026) -
Understanding and Mitigating Memorization in Generative Models via Sharpness of Probability Landscapes
by: Jeon, Dongjae, et al.
Published: (2024) -
Information-Theoretic Discrete Diffusion
by: Jeon, Moongyu, et al.
Published: (2025)