Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Solgi, Ryan, Madinei, Parsa, Tian, Jiayi, Swaminathan, Rupak, Liu, Jing, Susanj, Nathan, Zhang, Zheng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908579181625344
author Solgi, Ryan
Madinei, Parsa
Tian, Jiayi
Swaminathan, Rupak
Liu, Jing
Susanj, Nathan
Zhang, Zheng
author_facet Solgi, Ryan
Madinei, Parsa
Tian, Jiayi
Swaminathan, Rupak
Liu, Jing
Susanj, Nathan
Zhang, Zheng
contents Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and computing challenges in deployment. We present a novel low-rank compression framework to address this challenge. First, we upper bound the change of network loss via layer-wise activation-based compression errors, filling a theoretical gap in the literature. We then formulate low-rank model compression as a bi-objective optimization and prove that a single uniform tolerance yields surrogate Pareto-optimal heterogeneous ranks. Based on our theoretical insights, we propose Pareto-Guided Singular Value Decomposition (PGSVD), a zero-shot pipeline that improves activation-aware compression via Pareto-guided rank selection and alternating least-squares implementation. We apply PGSVD to both LLM and VLM, showing better accuracy at the same compression levels and inference speedup.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05544
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
Solgi, Ryan
Madinei, Parsa
Tian, Jiayi
Swaminathan, Rupak
Liu, Jing
Susanj, Nathan
Zhang, Zheng
Computation and Language
Machine Learning
Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and computing challenges in deployment. We present a novel low-rank compression framework to address this challenge. First, we upper bound the change of network loss via layer-wise activation-based compression errors, filling a theoretical gap in the literature. We then formulate low-rank model compression as a bi-objective optimization and prove that a single uniform tolerance yields surrogate Pareto-optimal heterogeneous ranks. Based on our theoretical insights, we propose Pareto-Guided Singular Value Decomposition (PGSVD), a zero-shot pipeline that improves activation-aware compression via Pareto-guided rank selection and alternating least-squares implementation. We apply PGSVD to both LLM and VLM, showing better accuracy at the same compression levels and inference speedup.
title Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.05544