Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Endo, Mark, Wang, Xiaohan, Yeung-Levy, Serena
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913968848633856
author Endo, Mark
Wang, Xiaohan
Yeung-Levy, Serena
author_facet Endo, Mark
Wang, Xiaohan
Yeung-Levy, Serena
contents Recent works on accelerating Vision-Language Models achieve strong performance across a variety of vision-language tasks despite highly compressing visual information. In this work, we examine the popular acceleration approach of early pruning of visual tokens inside the language model. Surprisingly, we find that while strong performance is maintained across many tasks, it exhibits drastically different behavior for a subset of vision-centric tasks such as localization. Upon further investigation, we uncover a core issue with the acceleration approach where most tokens towards the top of the image are pruned away. Yet, on many benchmarks aiming to evaluate vision-centric capabilities, strong performance persists with the flawed pruning strategy, highlighting these benchmarks' limited ability to assess fine-grained visual capabilities. Based on these findings, we propose FEATHER (Fast and Effective Acceleration wiTH Ensemble cRiteria), a straightforward approach that resolves the discovered early-layer pruning issue and further enhances the preservation of relevant tokens via multistage pruning with early uniform sampling to ensure broad image coverage. With comparable computational savings, we find that FEATHER achieves more than 5x performance improvement on the vision-centric localization benchmarks compared to the original acceleration approach.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13180
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
Endo, Mark
Wang, Xiaohan
Yeung-Levy, Serena
Computer Vision and Pattern Recognition
Recent works on accelerating Vision-Language Models achieve strong performance across a variety of vision-language tasks despite highly compressing visual information. In this work, we examine the popular acceleration approach of early pruning of visual tokens inside the language model. Surprisingly, we find that while strong performance is maintained across many tasks, it exhibits drastically different behavior for a subset of vision-centric tasks such as localization. Upon further investigation, we uncover a core issue with the acceleration approach where most tokens towards the top of the image are pruned away. Yet, on many benchmarks aiming to evaluate vision-centric capabilities, strong performance persists with the flawed pruning strategy, highlighting these benchmarks' limited ability to assess fine-grained visual capabilities. Based on these findings, we propose FEATHER (Fast and Effective Acceleration wiTH Ensemble cRiteria), a straightforward approach that resolves the discovered early-layer pruning issue and further enhances the preservation of relevant tokens via multistage pruning with early uniform sampling to ensure broad image coverage. With comparable computational savings, we find that FEATHER achieves more than 5x performance improvement on the vision-centric localization benchmarks compared to the original acceleration approach.
title Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.13180