Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Son, Jaemin, Choi, Sujin, Yun, Inyong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914366932123648
author Son, Jaemin
Choi, Sujin
Yun, Inyong
author_facet Son, Jaemin
Choi, Sujin
Yun, Inyong
contents Recent progress in vision-language models (VLMs) has led to impressive results in document understanding tasks, but their high computational demands remain a challenge. To mitigate the compute burdens, we propose a lightweight token pruning framework that filters out non-informative background regions from document images prior to VLM processing. A binary patch-level classifier removes non-text areas, and a max-pooling refinement step recovers fragmented text regions to enhance spatial coherence. Experiments on real-world document datasets demonstrate that our approach substantially lowers computational costs, while maintaining comparable accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06415
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models
Son, Jaemin
Choi, Sujin
Yun, Inyong
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Recent progress in vision-language models (VLMs) has led to impressive results in document understanding tasks, but their high computational demands remain a challenge. To mitigate the compute burdens, we propose a lightweight token pruning framework that filters out non-informative background regions from document images prior to VLM processing. A binary patch-level classifier removes non-text areas, and a max-pooling refinement step recovers fragmented text regions to enhance spatial coherence. Experiments on real-world document datasets demonstrate that our approach substantially lowers computational costs, while maintaining comparable accuracy.
title Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.06415