Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Zhaoqi, Zhang, Yingying, Li, Jian, Guo, Jianwei, Zhu, Qiannan, Huang, Hua
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909922320449536
author Xu, Zhaoqi
Zhang, Yingying
Li, Jian
Guo, Jianwei
Zhu, Qiannan
Huang, Hua
author_facet Xu, Zhaoqi
Zhang, Yingying
Li, Jian
Guo, Jianwei
Zhu, Qiannan
Huang, Hua
contents Recent advances in vision-language models (VLMs) have shown remarkable performance across multimodal tasks, yet their ever-growing scale poses severe challenges for deployment and efficiency. Existing compression methods often rely on heuristic importance metrics or empirical pruning rules, lacking theoretical guarantees about information preservation. In this work, we propose InfoPrune, an information-theoretic framework for adaptive structural compression of VLMs. Grounded in the Information Bottleneck principle, we formulate pruning as a trade-off between retaining task-relevant semantics and discarding redundant dependencies. To quantify the contribution of each attention head, we introduce an entropy-based effective rank (eRank) and employ the Kolmogorov--Smirnov (KS) distance to measure the divergence between original and compressed structures. This yields a unified criterion that jointly considers structural sparsity and informational efficiency. Building on this foundation, we further design two complementary schemes: (1) a training-based head pruning guided by the proposed information loss objective, and (2) a training-free FFN compression via adaptive low-rank approximation. Extensive experiments on VQAv2, TextVQA, and GQA demonstrate that InfoPrune achieves up to 3.2x FLOP reduction and 1.8x acceleration with negligible performance degradation, establishing a theoretically grounded and practically effective step toward efficient multimodal large models.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19518
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
Xu, Zhaoqi
Zhang, Yingying
Li, Jian
Guo, Jianwei
Zhu, Qiannan
Huang, Hua
Computer Vision and Pattern Recognition
Artificial Intelligence
Information Theory
Machine Learning
Recent advances in vision-language models (VLMs) have shown remarkable performance across multimodal tasks, yet their ever-growing scale poses severe challenges for deployment and efficiency. Existing compression methods often rely on heuristic importance metrics or empirical pruning rules, lacking theoretical guarantees about information preservation. In this work, we propose InfoPrune, an information-theoretic framework for adaptive structural compression of VLMs. Grounded in the Information Bottleneck principle, we formulate pruning as a trade-off between retaining task-relevant semantics and discarding redundant dependencies. To quantify the contribution of each attention head, we introduce an entropy-based effective rank (eRank) and employ the Kolmogorov--Smirnov (KS) distance to measure the divergence between original and compressed structures. This yields a unified criterion that jointly considers structural sparsity and informational efficiency. Building on this foundation, we further design two complementary schemes: (1) a training-based head pruning guided by the proposed information loss objective, and (2) a training-free FFN compression via adaptive low-rank approximation. Extensive experiments on VQAv2, TextVQA, and GQA demonstrate that InfoPrune achieves up to 3.2x FLOP reduction and 1.8x acceleration with negligible performance degradation, establishing a theoretically grounded and practically effective step toward efficient multimodal large models.
title Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Information Theory
Machine Learning
url https://arxiv.org/abs/2511.19518