CASP: Compression of Large Multimodal Models Based on Attention Sparsity

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gholami, Mohsen, Akbari, Mohammad, Cannons, Kevin, Zhang, Yong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915187639975936
author Gholami, Mohsen
Akbari, Mohammad
Cannons, Kevin
Zhang, Yong
author_facet Gholami, Mohsen
Akbari, Mohammad
Cannons, Kevin
Zhang, Yong
contents In this work, we propose an extreme compression technique for Large Multimodal Models (LMMs). While previous studies have explored quantization as an efficient post-training compression method for Large Language Models (LLMs), low-bit compression for multimodal models remains under-explored. The redundant nature of inputs in multimodal models results in a highly sparse attention matrix. We theoretically and experimentally demonstrate that the attention matrix's sparsity bounds the compression error of the Query and Key weight matrices. Based on this, we introduce CASP, a model compression technique for LMMs. Our approach performs a data-aware low-rank decomposition on the Query and Key weight matrix, followed by quantization across all layers based on an optimal bit allocation process. CASP is compatible with any quantization technique and enhances state-of-the-art 2-bit quantization methods (AQLM and QuIP#) by an average of 21% on image- and video-language benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_05936
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CASP: Compression of Large Multimodal Models Based on Attention Sparsity
Gholami, Mohsen
Akbari, Mohammad
Cannons, Kevin
Zhang, Yong
Computer Vision and Pattern Recognition
In this work, we propose an extreme compression technique for Large Multimodal Models (LMMs). While previous studies have explored quantization as an efficient post-training compression method for Large Language Models (LLMs), low-bit compression for multimodal models remains under-explored. The redundant nature of inputs in multimodal models results in a highly sparse attention matrix. We theoretically and experimentally demonstrate that the attention matrix's sparsity bounds the compression error of the Query and Key weight matrices. Based on this, we introduce CASP, a model compression technique for LMMs. Our approach performs a data-aware low-rank decomposition on the Query and Key weight matrix, followed by quantization across all layers based on an optimal bit allocation process. CASP is compatible with any quantization technique and enhances state-of-the-art 2-bit quantization methods (AQLM and QuIP#) by an average of 21% on image- and video-language benchmarks.
title CASP: Compression of Large Multimodal Models Based on Attention Sparsity
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.05936