FoPru: Focal Pruning for Efficient Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Lei, Huang, Weizhe, Liu, Tongxuan, Zeng, Yuting, Li, Jing, Cheng, Lechao, Xu, Xiaohua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909398302982144
author Jiang, Lei
Huang, Weizhe
Liu, Tongxuan
Zeng, Yuting
Li, Jing
Cheng, Lechao
Xu, Xiaohua
author_facet Jiang, Lei
Huang, Weizhe
Liu, Tongxuan
Zeng, Yuting
Li, Jing
Cheng, Lechao
Xu, Xiaohua
contents Large Vision-Language Models (LVLMs) represent a significant advancement toward achieving superior multimodal capabilities by enabling powerful Large Language Models (LLMs) to understand visual input. Typically, LVLMs utilize visual encoders, such as CLIP, to transform images into visual tokens, which are then aligned with textual tokens through projection layers before being input into the LLM for inference. Although existing LVLMs have achieved significant success, their inference efficiency is still limited by the substantial number of visual tokens and the potential redundancy among them. To mitigate this issue, we propose Focal Pruning (FoPru), a training-free method that prunes visual tokens based on the attention-based token significance derived from the vision encoder. Specifically, we introduce two alternative pruning strategies: 1) the rank strategy, which leverages all token significance scores to retain more critical tokens in a global view; 2) the row strategy, which focuses on preserving continuous key information in images from a local perspective. Finally, the selected tokens are reordered to maintain their original positional relationships. Extensive experiments across various LVLMs and multimodal datasets demonstrate that our method can prune a large number of redundant tokens while maintaining high accuracy, leading to significant improvements in inference efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14164
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FoPru: Focal Pruning for Efficient Large Vision-Language Models
Jiang, Lei
Huang, Weizhe
Liu, Tongxuan
Zeng, Yuting
Li, Jing
Cheng, Lechao
Xu, Xiaohua
Computer Vision and Pattern Recognition
Artificial Intelligence
Large Vision-Language Models (LVLMs) represent a significant advancement toward achieving superior multimodal capabilities by enabling powerful Large Language Models (LLMs) to understand visual input. Typically, LVLMs utilize visual encoders, such as CLIP, to transform images into visual tokens, which are then aligned with textual tokens through projection layers before being input into the LLM for inference. Although existing LVLMs have achieved significant success, their inference efficiency is still limited by the substantial number of visual tokens and the potential redundancy among them. To mitigate this issue, we propose Focal Pruning (FoPru), a training-free method that prunes visual tokens based on the attention-based token significance derived from the vision encoder. Specifically, we introduce two alternative pruning strategies: 1) the rank strategy, which leverages all token significance scores to retain more critical tokens in a global view; 2) the row strategy, which focuses on preserving continuous key information in images from a local perspective. Finally, the selected tokens are reordered to maintain their original positional relationships. Extensive experiments across various LVLMs and multimodal datasets demonstrate that our method can prune a large number of redundant tokens while maintaining high accuracy, leading to significant improvements in inference efficiency.
title FoPru: Focal Pruning for Efficient Large Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.14164