AdaSVD: Adaptive Singular Value Decomposition for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhiteng, Xia, Mingyuan, Zhang, Jingyuan, Hui, Zheng, Qin, Haotong, Kong, Linghe, Zhang, Yulun, Yang, Xiaokang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914054963986432
author Li, Zhiteng
Xia, Mingyuan
Zhang, Jingyuan
Hui, Zheng
Qin, Haotong
Kong, Linghe
Zhang, Yulun
Yang, Xiaokang
author_facet Li, Zhiteng
Xia, Mingyuan
Zhang, Jingyuan
Hui, Zheng
Qin, Haotong
Kong, Linghe
Zhang, Yulun
Yang, Xiaokang
contents Large language models (LLMs) have achieved remarkable success in natural language processing (NLP) tasks, yet their substantial memory requirements present significant challenges for deployment on resource-constrained devices. Singular Value Decomposition (SVD) has emerged as a promising compression technique for LLMs, offering considerable reductions in memory overhead. However, existing SVD-based methods often struggle to effectively mitigate the errors introduced by SVD truncation, leading to a noticeable performance gap when compared to the original models. Furthermore, applying a uniform compression ratio across all transformer layers fails to account for the varying importance of different layers. To address these challenges, we propose AdaSVD, an adaptive SVD-based LLM compression approach. Specifically, AdaSVD introduces adaComp, which adaptively compensates for SVD truncation errors by alternately updating the singular matrices $\mathcal{U}$ and $\mathcal{V}^\top$. Additionally, AdaSVD introduces adaCR, which adaptively assigns layer-specific compression ratios based on the relative importance of each layer. Extensive experiments across multiple LLM/VLM families and evaluation metrics demonstrate that AdaSVD consistently outperforms state-of-the-art (SOTA) SVD-based methods, achieving superior performance with significantly reduced memory requirements. Code and models of AdaSVD will be available at https://github.com/ZHITENGLI/AdaSVD.
format Preprint
id arxiv_https___arxiv_org_abs_2502_01403
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdaSVD: Adaptive Singular Value Decomposition for Large Language Models
Li, Zhiteng
Xia, Mingyuan
Zhang, Jingyuan
Hui, Zheng
Qin, Haotong
Kong, Linghe
Zhang, Yulun
Yang, Xiaokang
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Large language models (LLMs) have achieved remarkable success in natural language processing (NLP) tasks, yet their substantial memory requirements present significant challenges for deployment on resource-constrained devices. Singular Value Decomposition (SVD) has emerged as a promising compression technique for LLMs, offering considerable reductions in memory overhead. However, existing SVD-based methods often struggle to effectively mitigate the errors introduced by SVD truncation, leading to a noticeable performance gap when compared to the original models. Furthermore, applying a uniform compression ratio across all transformer layers fails to account for the varying importance of different layers. To address these challenges, we propose AdaSVD, an adaptive SVD-based LLM compression approach. Specifically, AdaSVD introduces adaComp, which adaptively compensates for SVD truncation errors by alternately updating the singular matrices $\mathcal{U}$ and $\mathcal{V}^\top$. Additionally, AdaSVD introduces adaCR, which adaptively assigns layer-specific compression ratios based on the relative importance of each layer. Extensive experiments across multiple LLM/VLM families and evaluation metrics demonstrate that AdaSVD consistently outperforms state-of-the-art (SOTA) SVD-based methods, achieving superior performance with significantly reduced memory requirements. Code and models of AdaSVD will be available at https://github.com/ZHITENGLI/AdaSVD.
title AdaSVD: Adaptive Singular Value Decomposition for Large Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.01403