Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cho, Yoonjun, Kim, Soeun, Jeon, Dongjae, Lee, Kyelim, Lee, Beomsoo, No, Albert
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908391037730816
author Cho, Yoonjun
Kim, Soeun
Jeon, Dongjae
Lee, Kyelim
Lee, Beomsoo
No, Albert
author_facet Cho, Yoonjun
Kim, Soeun
Jeon, Dongjae
Lee, Kyelim
Lee, Beomsoo
No, Albert
contents Decomposing weight matrices into quantization and low-rank components ($\mathbf{W} \approx \mathbf{Q} + \mathbf{L}\mathbf{R}$) is a widely used technique for compressing large language models (LLMs). Existing joint optimization methods iteratively alternate between quantization and low-rank approximation. However, these methods tend to prioritize one component at the expense of the other, resulting in suboptimal decompositions that fail to leverage each component's unique strengths. In this work, we introduce Outlier-Driven Low-Rank Initialization (ODLRI), which assigns low-rank components the specific role of capturing activation-sensitive weights. This structured decomposition mitigates outliers' negative impact on quantization, enabling more effective balance between quantization and low-rank approximation. Experiments on Llama2 (7B, 13B, 70B), Llama3-8B, and Mistral-7B demonstrate that incorporating ODLRI into the joint optimization framework consistently reduces activation-aware error, minimizes quantization scale, and improves perplexity and zero-shot accuracy in low-bit settings.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02077
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
Cho, Yoonjun
Kim, Soeun
Jeon, Dongjae
Lee, Kyelim
Lee, Beomsoo
No, Albert
Machine Learning
Artificial Intelligence
Computation and Language
Decomposing weight matrices into quantization and low-rank components ($\mathbf{W} \approx \mathbf{Q} + \mathbf{L}\mathbf{R}$) is a widely used technique for compressing large language models (LLMs). Existing joint optimization methods iteratively alternate between quantization and low-rank approximation. However, these methods tend to prioritize one component at the expense of the other, resulting in suboptimal decompositions that fail to leverage each component's unique strengths. In this work, we introduce Outlier-Driven Low-Rank Initialization (ODLRI), which assigns low-rank components the specific role of capturing activation-sensitive weights. This structured decomposition mitigates outliers' negative impact on quantization, enabling more effective balance between quantization and low-rank approximation. Experiments on Llama2 (7B, 13B, 70B), Llama3-8B, and Mistral-7B demonstrate that incorporating ODLRI into the joint optimization framework consistently reduces activation-aware error, minimizes quantization scale, and improves perplexity and zero-shot accuracy in low-bit settings.
title Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.02077