Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abbasi, Ali, Thrash, Chayne, Qin, Haoran, Sharma, Shansita, Seifi, Sepehr, Kolouri, Soheil
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908807890731008
author Abbasi, Ali
Thrash, Chayne
Qin, Haoran
Sharma, Shansita
Seifi, Sepehr
Kolouri, Soheil
author_facet Abbasi, Ali
Thrash, Chayne
Qin, Haoran
Sharma, Shansita
Seifi, Sepehr
Kolouri, Soheil
contents Advances in large language models have driven strong performance across many tasks, but their memory and compute costs still hinder deployment. SVD-based compression reduces storage and can speed up inference via low-rank factors, yet performance depends on how rank is allocated under a global compression ratio. Prior methods often use homogeneous ranks for similarly sized matrices, despite large differences in loss sensitivity, or rely on expensive iterative pre-truncation optimization to determine per matrix ranks. We propose \textbf{Zero Sum SVD} (\textbf{ZS-SVD}), a post-training method that performs \emph{global} singular component selection using activation whitening and first-order calibration loss estimates in whitened coordinates. \textbf{ZS-SVD} prunes components across the whole model with a \textbf{zero sum} rule that keeps the cumulative predicted loss change near zero, automatically yielding heterogeneous ranks without solving a rank allocation optimization. Motivated by evidence that gradients near pretrained solutions exhibit low rank structure, we also introduce an optional lightweight correction that applies a \textbf{single} projected gradient update after truncation, followed by re-truncation. Extensive experiments across multiple LLM architectures show consistent gains across diverse benchmarks and compression ratios. Code is available at https://github.com/mint-vu/Zero-Sum-SVD
format Preprint
id arxiv_https___arxiv_org_abs_2602_02848
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
Abbasi, Ali
Thrash, Chayne
Qin, Haoran
Sharma, Shansita
Seifi, Sepehr
Kolouri, Soheil
Machine Learning
Advances in large language models have driven strong performance across many tasks, but their memory and compute costs still hinder deployment. SVD-based compression reduces storage and can speed up inference via low-rank factors, yet performance depends on how rank is allocated under a global compression ratio. Prior methods often use homogeneous ranks for similarly sized matrices, despite large differences in loss sensitivity, or rely on expensive iterative pre-truncation optimization to determine per matrix ranks. We propose \textbf{Zero Sum SVD} (\textbf{ZS-SVD}), a post-training method that performs \emph{global} singular component selection using activation whitening and first-order calibration loss estimates in whitened coordinates. \textbf{ZS-SVD} prunes components across the whole model with a \textbf{zero sum} rule that keeps the cumulative predicted loss change near zero, automatically yielding heterogeneous ranks without solving a rank allocation optimization. Motivated by evidence that gradients near pretrained solutions exhibit low rank structure, we also introduce an optional lightweight correction that applies a \textbf{single} projected gradient update after truncation, followed by re-truncation. Extensive experiments across multiple LLM architectures show consistent gains across diverse benchmarks and compression ratios. Code is available at https://github.com/mint-vu/Zero-Sum-SVD
title Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
topic Machine Learning
url https://arxiv.org/abs/2602.02848