SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866912409879314432 |
|---|---|
| author | Zhang, Yuan Fan, Chun-Kai Ma, Junpeng Zheng, Wenzhao Huang, Tao Cheng, Kuan Gudovskiy, Denis Okuno, Tomoyuki Nakata, Yohei Keutzer, Kurt Zhang, Shanghang |
| author_facet | Zhang, Yuan Fan, Chun-Kai Ma, Junpeng Zheng, Wenzhao Huang, Tao Cheng, Kuan Gudovskiy, Denis Okuno, Tomoyuki Nakata, Yohei Keutzer, Kurt Zhang, Shanghang |
| contents | In vision-language models (VLMs), visual tokens usually bear a significant amount of computational overhead despite sparsity of information in them when compared to text tokens. To address this, most existing methods learn a network to prune redundant visual tokens using certain training data. Differently, we propose a text-guided training-free token optimization mechanism dubbed SparseVLM that eliminates the need of extra parameters or fine-tuning costs. Given that visual tokens complement text tokens in VLM's linguistic reasoning, we select relevant text tokens to rate the significance of visual tokens using self-attention matrices and, then, prune visual tokens using the proposed strategy to maximize sparsity while retaining information. In particular, we introduce a rank-based strategy to adaptively determine the sparsification ratio for each layer, alongside a token recycling method that compresses pruned tokens into more compact representations. Experimental results show that SparseVLM increases the efficiency of various VLMs in a number of image and video understanding tasks. For example, LLaVA when equipped with SparseVLM achieves 54% reduction in FLOPs, 37% decrease in CUDA latency while maintaining 97% of its original accuracy. Our code is available at https://github.com/Gumpest/SparseVLMs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_04417 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference Zhang, Yuan Fan, Chun-Kai Ma, Junpeng Zheng, Wenzhao Huang, Tao Cheng, Kuan Gudovskiy, Denis Okuno, Tomoyuki Nakata, Yohei Keutzer, Kurt Zhang, Shanghang Computer Vision and Pattern Recognition In vision-language models (VLMs), visual tokens usually bear a significant amount of computational overhead despite sparsity of information in them when compared to text tokens. To address this, most existing methods learn a network to prune redundant visual tokens using certain training data. Differently, we propose a text-guided training-free token optimization mechanism dubbed SparseVLM that eliminates the need of extra parameters or fine-tuning costs. Given that visual tokens complement text tokens in VLM's linguistic reasoning, we select relevant text tokens to rate the significance of visual tokens using self-attention matrices and, then, prune visual tokens using the proposed strategy to maximize sparsity while retaining information. In particular, we introduce a rank-based strategy to adaptively determine the sparsification ratio for each layer, alongside a token recycling method that compresses pruned tokens into more compact representations. Experimental results show that SparseVLM increases the efficiency of various VLMs in a number of image and video understanding tasks. For example, LLaVA when equipped with SparseVLM achieves 54% reduction in FLOPs, 37% decrease in CUDA latency while maintaining 97% of its original accuracy. Our code is available at https://github.com/Gumpest/SparseVLMs. |
| title | SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2410.04417 |