CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Kangjie, Huang, Wenxuan, Zhou, Xin, Zhou, Boxiang, Song, Dejia, Xie, Yuan, Zhang, Baochang, Ma, Lizhuang, Chen, Nemo, Tang, Xu, Hu, Yao, Lin, Shaohui
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908816392585216
author Zhang, Kangjie
Huang, Wenxuan
Zhou, Xin
Zhou, Boxiang
Song, Dejia
Xie, Yuan
Zhang, Baochang
Ma, Lizhuang
Chen, Nemo
Tang, Xu
Hu, Yao
Lin, Shaohui
author_facet Zhang, Kangjie
Huang, Wenxuan
Zhou, Xin
Zhou, Boxiang
Song, Dejia
Xie, Yuan
Zhang, Baochang
Ma, Lizhuang
Chen, Nemo
Tang, Xu
Hu, Yao
Lin, Shaohui
contents Contrastive Language-Image Pre-training (CLIP) has achieved widely applications in various computer vision tasks, e.g., text-to-image generation, Image-Text retrieval and Image captioning. However, CLIP suffers from high memory and computation cost, which prohibits its usage to the resource-limited application scenarios. Existing CLIP compression methods typically reduce the size of pre-trained CLIP weights by selecting their subset as weight inheritance for further retraining via mask optimization or important weight measurement. However, these select-based weight inheritance often compromises the feature presentation ability, especially on the extreme compression. In this paper, we propose a novel mapping-based CLIP compression framework, CLIP-Map. It leverages learnable matrices to map and combine pretrained weights by Full-Mapping with Kronecker Factorization, aiming to preserve as much information from the original weights as possible. To mitigate the optimization challenges introduced by the learnable mapping, we propose Diagonal Inheritance Initialization to reduce the distribution shifting problem for efficient and effective mapping learning. Extensive experimental results demonstrate that the proposed CLIP-Map outperforms select-based frameworks across various compression ratios, with particularly significant gains observed under high compression settings.
format Preprint
id arxiv_https___arxiv_org_abs_2602_05909
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
Zhang, Kangjie
Huang, Wenxuan
Zhou, Xin
Zhou, Boxiang
Song, Dejia
Xie, Yuan
Zhang, Baochang
Ma, Lizhuang
Chen, Nemo
Tang, Xu
Hu, Yao
Lin, Shaohui
Computer Vision and Pattern Recognition
Contrastive Language-Image Pre-training (CLIP) has achieved widely applications in various computer vision tasks, e.g., text-to-image generation, Image-Text retrieval and Image captioning. However, CLIP suffers from high memory and computation cost, which prohibits its usage to the resource-limited application scenarios. Existing CLIP compression methods typically reduce the size of pre-trained CLIP weights by selecting their subset as weight inheritance for further retraining via mask optimization or important weight measurement. However, these select-based weight inheritance often compromises the feature presentation ability, especially on the extreme compression. In this paper, we propose a novel mapping-based CLIP compression framework, CLIP-Map. It leverages learnable matrices to map and combine pretrained weights by Full-Mapping with Kronecker Factorization, aiming to preserve as much information from the original weights as possible. To mitigate the optimization challenges introduced by the learnable mapping, we propose Diagonal Inheritance Initialization to reduce the distribution shifting problem for efficient and effective mapping learning. Extensive experimental results demonstrate that the proposed CLIP-Map outperforms select-based frameworks across various compression ratios, with particularly significant gains observed under high compression settings.
title CLIP-Map: Structured Matrix Mapping for Parameter-Efficient CLIP Compression
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.05909