TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Shu-Hao, Tang, Wei-Cheng, Wu, Chen, Hu, Peng, Li, Nan, Zhang, Liang-Jie, Zhang, Qi, Zhang, Shao-Qun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915576253775872
author Zhang, Shu-Hao
Tang, Wei-Cheng
Wu, Chen
Hu, Peng
Li, Nan
Zhang, Liang-Jie
Zhang, Qi
Zhang, Shao-Qun
author_facet Zhang, Shu-Hao
Tang, Wei-Cheng
Wu, Chen
Hu, Peng
Li, Nan
Zhang, Liang-Jie
Zhang, Qi
Zhang, Shao-Qun
contents Recent years have witnessed an increasing interest in image-text contrastive modeling, exemplified by models such as Contrastive Language-Image Pretraining (CLIP). In this paper, we propose the TernaryCLIP, a lightweight computational framework that converts connection weights of both vision and text encoders of CLIP into the ternary format, instead of full-precision or floating ones. TernaryCLIP incorporates quantization-aware training and distillation modules, preventing precision degradation and enabling low-cost and high-efficiency computations. Comprehensive experiments demonstrate that TernaryCLIP can achieve up to 99\% ternarized weights with 1.58-bit representation, 16.98 $\times$ compression ratio, 2.3 $\times$ inference acceleration, 16 $\times$ storage reduction, 10 $\times$ memory optimization, and 60\% sparsity while maintaining promising performance on zero-shot image classification and image-text retrieval tasks across 41 commonly used datasets. Our work highlights the feasibility of extreme quantization for large multimodal models, supporting effective and efficient deployment on resource-constrained devices. The model and code can be accessed from Hugging Face and GitHub.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21879
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
Zhang, Shu-Hao
Tang, Wei-Cheng
Wu, Chen
Hu, Peng
Li, Nan
Zhang, Liang-Jie
Zhang, Qi
Zhang, Shao-Qun
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent years have witnessed an increasing interest in image-text contrastive modeling, exemplified by models such as Contrastive Language-Image Pretraining (CLIP). In this paper, we propose the TernaryCLIP, a lightweight computational framework that converts connection weights of both vision and text encoders of CLIP into the ternary format, instead of full-precision or floating ones. TernaryCLIP incorporates quantization-aware training and distillation modules, preventing precision degradation and enabling low-cost and high-efficiency computations. Comprehensive experiments demonstrate that TernaryCLIP can achieve up to 99\% ternarized weights with 1.58-bit representation, 16.98 $\times$ compression ratio, 2.3 $\times$ inference acceleration, 16 $\times$ storage reduction, 10 $\times$ memory optimization, and 60\% sparsity while maintaining promising performance on zero-shot image classification and image-text retrieval tasks across 41 commonly used datasets. Our work highlights the feasibility of extreme quantization for large multimodal models, supporting effective and efficient deployment on resource-constrained devices. The model and code can be accessed from Hugging Face and GitHub.
title TernaryCLIP: Efficiently Compressing Vision-Language Models with Ternary Weights and Distilled Knowledge
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.21879