EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zekun, Ma, Minghua, Wang, Zexin, Mu, Rongchuan, Shan, Liping, Liu, Ming, Qin, Bing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916770516828160
author Wang, Zekun
Ma, Minghua
Wang, Zexin
Mu, Rongchuan
Shan, Liping
Liu, Ming
Qin, Bing
author_facet Wang, Zekun
Ma, Minghua
Wang, Zexin
Mu, Rongchuan
Shan, Liping
Liu, Ming
Qin, Bing
contents Large Vision-Language Models (LVLMs) have achieved remarkable success, yet their significant computational demands hinder practical deployment. While efforts to improve LVLM efficiency are growing, existing methods lack comprehensive evaluation across diverse backbones, benchmarks, and metrics. In this work, we systematically evaluate mainstream acceleration techniques for LVLMs, categorized into token and parameter compression. We introduce EffiVLM-Bench, a unified framework for assessing not only absolute performance but also generalization and loyalty, while exploring Pareto-optimal trade-offs. Our extensive experiments and in-depth analyses offer insights into optimal strategies for accelerating LVLMs. We open-source code and recipes for EffiVLM-Bench to foster future research.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00479
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
Wang, Zekun
Ma, Minghua
Wang, Zexin
Mu, Rongchuan
Shan, Liping
Liu, Ming
Qin, Bing
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
Large Vision-Language Models (LVLMs) have achieved remarkable success, yet their significant computational demands hinder practical deployment. While efforts to improve LVLM efficiency are growing, existing methods lack comprehensive evaluation across diverse backbones, benchmarks, and metrics. In this work, we systematically evaluate mainstream acceleration techniques for LVLMs, categorized into token and parameter compression. We introduce EffiVLM-Bench, a unified framework for assessing not only absolute performance but also generalization and loyalty, while exploring Pareto-optimal trade-offs. Our extensive experiments and in-depth analyses offer insights into optimal strategies for accelerating LVLMs. We open-source code and recipes for EffiVLM-Bench to foster future research.
title EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
topic Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2506.00479