Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914366932123648 |
|---|---|
| author | Son, Jaemin Choi, Sujin Yun, Inyong |
| author_facet | Son, Jaemin Choi, Sujin Yun, Inyong |
| contents | Recent progress in vision-language models (VLMs) has led to impressive results in document understanding tasks, but their high computational demands remain a challenge. To mitigate the compute burdens, we propose a lightweight token pruning framework that filters out non-informative background regions from document images prior to VLM processing. A binary patch-level classifier removes non-text areas, and a max-pooling refinement step recovers fragmented text regions to enhance spatial coherence. Experiments on real-world document datasets demonstrate that our approach substantially lowers computational costs, while maintaining comparable accuracy. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_06415 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models Son, Jaemin Choi, Sujin Yun, Inyong Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Recent progress in vision-language models (VLMs) has led to impressive results in document understanding tasks, but their high computational demands remain a challenge. To mitigate the compute burdens, we propose a lightweight token pruning framework that filters out non-informative background regions from document images prior to VLM processing. A binary patch-level classifier removes non-text areas, and a max-pooling refinement step recovers fragmented text regions to enhance spatial coherence. Experiments on real-world document datasets demonstrate that our approach substantially lowers computational costs, while maintaining comparable accuracy. |
| title | Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2509.06415 |