Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908579181625344 |
|---|---|
| author | Solgi, Ryan Madinei, Parsa Tian, Jiayi Swaminathan, Rupak Liu, Jing Susanj, Nathan Zhang, Zheng |
| author_facet | Solgi, Ryan Madinei, Parsa Tian, Jiayi Swaminathan, Rupak Liu, Jing Susanj, Nathan Zhang, Zheng |
| contents | Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and computing challenges in deployment. We present a novel low-rank compression framework to address this challenge. First, we upper bound the change of network loss via layer-wise activation-based compression errors, filling a theoretical gap in the literature. We then formulate low-rank model compression as a bi-objective optimization and prove that a single uniform tolerance yields surrogate Pareto-optimal heterogeneous ranks. Based on our theoretical insights, we propose Pareto-Guided Singular Value Decomposition (PGSVD), a zero-shot pipeline that improves activation-aware compression via Pareto-guided rank selection and alternating least-squares implementation. We apply PGSVD to both LLM and VLM, showing better accuracy at the same compression levels and inference speedup. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_05544 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM Solgi, Ryan Madinei, Parsa Tian, Jiayi Swaminathan, Rupak Liu, Jing Susanj, Nathan Zhang, Zheng Computation and Language Machine Learning Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and computing challenges in deployment. We present a novel low-rank compression framework to address this challenge. First, we upper bound the change of network loss via layer-wise activation-based compression errors, filling a theoretical gap in the literature. We then formulate low-rank model compression as a bi-objective optimization and prove that a single uniform tolerance yields surrogate Pareto-optimal heterogeneous ranks. Based on our theoretical insights, we propose Pareto-Guided Singular Value Decomposition (PGSVD), a zero-shot pipeline that improves activation-aware compression via Pareto-guided rank selection and alternating least-squares implementation. We apply PGSVD to both LLM and VLM, showing better accuracy at the same compression levels and inference speedup. |
| title | Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2510.05544 |