Rethinking Pruning for Vision-Language Models: Strategies for Effective Sparsity and Performance Restoration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Shwai, Li, Ang, Chen, Tianlong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911930824785920
author He, Shwai
Li, Ang
Chen, Tianlong
author_facet He, Shwai
Li, Ang
Chen, Tianlong
contents Vision-Language Models (VLMs) integrate information from multiple modalities and have shown remarkable success across various tasks. However, deploying large-scale VLMs in resource-constrained scenarios is challenging. Pruning followed by finetuning offers a potential solution but remains underexplored for VLMs. This study addresses two key questions: how to distribute sparsity across different modality-specific models, and how to restore the performance of pruned sparse VLMs. Our preliminary studies identified two effective pruning settings: applying the same sparsity to both vision and language models, and pruning only the language models. While LoRA finetuning aims to restore sparse models, it faces challenges due to incompatibility with sparse models, disrupting the pruned sparsity. To overcome these issues, we propose SparseLoRA, which applies sparsity directly to LoRA weights. Our experimental results demonstrate significant improvements, including an 11.3\% boost under 2:4 sparsity and a 47.6\% enhancement under unstructured 70\% sparsity. Code is released at: \url{https://github.com/Shwai-He/VLM-Compression}.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02424
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rethinking Pruning for Vision-Language Models: Strategies for Effective Sparsity and Performance Restoration
He, Shwai
Li, Ang
Chen, Tianlong
Machine Learning
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs) integrate information from multiple modalities and have shown remarkable success across various tasks. However, deploying large-scale VLMs in resource-constrained scenarios is challenging. Pruning followed by finetuning offers a potential solution but remains underexplored for VLMs. This study addresses two key questions: how to distribute sparsity across different modality-specific models, and how to restore the performance of pruned sparse VLMs. Our preliminary studies identified two effective pruning settings: applying the same sparsity to both vision and language models, and pruning only the language models. While LoRA finetuning aims to restore sparse models, it faces challenges due to incompatibility with sparse models, disrupting the pruned sparsity. To overcome these issues, we propose SparseLoRA, which applies sparsity directly to LoRA weights. Our experimental results demonstrate significant improvements, including an 11.3\% boost under 2:4 sparsity and a 47.6\% enhancement under unstructured 70\% sparsity. Code is released at: \url{https://github.com/Shwai-He/VLM-Compression}.
title Rethinking Pruning for Vision-Language Models: Strategies for Effective Sparsity and Performance Restoration
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.02424