LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lv, Chengtao, Zhang, Bilang, Yong, Yang, Gong, Ruihao, Huang, Yushi, Gu, Shiqiao, Wu, Jiajun, Shi, Yumeng, Guo, Jinyang, Wang, Wenya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917085271031808
author Lv, Chengtao
Zhang, Bilang
Yong, Yang
Gong, Ruihao
Huang, Yushi
Gu, Shiqiao
Wu, Jiajun
Shi, Yumeng
Guo, Jinyang
Wang, Wenya
author_facet Lv, Chengtao
Zhang, Bilang
Yong, Yang
Gong, Ruihao
Huang, Yushi
Gu, Shiqiao
Wu, Jiajun
Shi, Yumeng
Guo, Jinyang
Wang, Wenya
contents Large Vision-Language Models (VLMs) exhibit impressive multi-modal capabilities but suffer from prohibitive computational and memory demands, due to their long visual token sequences and massive parameter sizes. To address these issues, recent works have proposed training-free compression methods. However, existing efforts often suffer from three major limitations: (1) Current approaches do not decompose techniques into comparable modules, hindering fair evaluation across spatial and temporal redundancy. (2) Evaluation confined to simple single-turn tasks, failing to reflect performance in realistic scenarios. (3) Isolated use of individual compression techniques, without exploring their joint potential. To overcome these gaps, we introduce LLMC+, a comprehensive VLM compression benchmark with a versatile, plug-and-play toolkit. LLMC+ supports over 20 algorithms across five representative VLM families and enables systematic study of token-level and model-level compression. Our benchmark reveals that: (1) Spatial and temporal redundancies demand distinct technical strategies. (2) Token reduction methods degrade significantly in multi-turn dialogue and detail-sensitive tasks. (3) Combining token and model compression achieves extreme compression with minimal performance loss. We believe LLMC+ will facilitate fair evaluation and inspire future research in efficient VLM. Our code is available at https://github.com/ModelTC/LightCompress.
format Preprint
id arxiv_https___arxiv_org_abs_2508_09981
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
Lv, Chengtao
Zhang, Bilang
Yong, Yang
Gong, Ruihao
Huang, Yushi
Gu, Shiqiao
Wu, Jiajun
Shi, Yumeng
Guo, Jinyang
Wang, Wenya
Computer Vision and Pattern Recognition
Large Vision-Language Models (VLMs) exhibit impressive multi-modal capabilities but suffer from prohibitive computational and memory demands, due to their long visual token sequences and massive parameter sizes. To address these issues, recent works have proposed training-free compression methods. However, existing efforts often suffer from three major limitations: (1) Current approaches do not decompose techniques into comparable modules, hindering fair evaluation across spatial and temporal redundancy. (2) Evaluation confined to simple single-turn tasks, failing to reflect performance in realistic scenarios. (3) Isolated use of individual compression techniques, without exploring their joint potential. To overcome these gaps, we introduce LLMC+, a comprehensive VLM compression benchmark with a versatile, plug-and-play toolkit. LLMC+ supports over 20 algorithms across five representative VLM families and enables systematic study of token-level and model-level compression. Our benchmark reveals that: (1) Spatial and temporal redundancies demand distinct technical strategies. (2) Token reduction methods degrade significantly in multi-turn dialogue and detail-sensitive tasks. (3) Combining token and model compression achieves extreme compression with minimal performance loss. We believe LLMC+ will facilitate fair evaluation and inspire future research in efficient VLM. Our code is available at https://github.com/ModelTC/LightCompress.
title LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.09981