Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Qinsi, Ke, Jinghan, Tomizuka, Masayoshi, Chen, Yiran, Keutzer, Kurt, Xu, Chenfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929698700787712
author Wang, Qinsi
Ke, Jinghan
Tomizuka, Masayoshi
Chen, Yiran
Keutzer, Kurt
Xu, Chenfeng
author_facet Wang, Qinsi
Ke, Jinghan
Tomizuka, Masayoshi
Chen, Yiran
Keutzer, Kurt
Xu, Chenfeng
contents We provide a new LLM-compression solution via SVD, unlocking new possibilities for LLM compression beyond quantization and pruning. We point out that the optimal use of SVD lies in truncating activations, rather than merely using activations as an optimization distance. Building on this principle, we address three critical challenges in SVD-based LLM compression: including (1) How can we determine the optimal activation truncation position for each weight matrix in LLMs? (2) How can we efficiently reconstruct the weight matrices based on truncated activations? (3) How can we address the inherent "injection" nature that results in the information loss of the SVD? We propose Dobi-SVD, which establishes a new, principled approach to SVD-based LLM compression.
format Preprint
id arxiv_https___arxiv_org_abs_2502_02723
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives
Wang, Qinsi
Ke, Jinghan
Tomizuka, Masayoshi
Chen, Yiran
Keutzer, Kurt
Xu, Chenfeng
Machine Learning
We provide a new LLM-compression solution via SVD, unlocking new possibilities for LLM compression beyond quantization and pruning. We point out that the optimal use of SVD lies in truncating activations, rather than merely using activations as an optimization distance. Building on this principle, we address three critical challenges in SVD-based LLM compression: including (1) How can we determine the optimal activation truncation position for each weight matrix in LLMs? (2) How can we efficiently reconstruct the weight matrices based on truncated activations? (3) How can we address the inherent "injection" nature that results in the information loss of the SVD? We propose Dobi-SVD, which establishes a new, principled approach to SVD-based LLM compression.
title Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives
topic Machine Learning
url https://arxiv.org/abs/2502.02723