DocVLM: Make Your VLM an Efficient Reader

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nacson, Mor Shpigel, Aberdam, Aviad, Ganz, Roy, Avraham, Elad Ben, Golts, Alona, Kittenplon, Yair, Mazor, Shai, Litman, Ron
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915060651130880
author Nacson, Mor Shpigel
Aberdam, Aviad
Ganz, Roy
Avraham, Elad Ben
Golts, Alona
Kittenplon, Yair
Mazor, Shai
Litman, Ron
author_facet Nacson, Mor Shpigel
Aberdam, Aviad
Ganz, Roy
Avraham, Elad Ben
Golts, Alona
Kittenplon, Yair
Mazor, Shai
Litman, Ron
contents Vision-Language Models (VLMs) excel in diverse visual tasks but face challenges in document understanding, which requires fine-grained text processing. While typical visual tasks perform well with low-resolution inputs, reading-intensive applications demand high-resolution, resulting in significant computational overhead. Using OCR-extracted text in VLM prompts partially addresses this issue but underperforms compared to full-resolution counterpart, as it lacks the complete visual context needed for optimal performance. We introduce DocVLM, a method that integrates an OCR-based modality into VLMs to enhance document processing while preserving original weights. Our approach employs an OCR encoder to capture textual content and layout, compressing these into a compact set of learned queries incorporated into the VLM. Comprehensive evaluations across leading VLMs show that DocVLM significantly reduces reliance on high-resolution images for document understanding. In limited-token regimes (448$\times$448), DocVLM with 64 learned queries improves DocVQA results from 56.0% to 86.6% when integrated with InternVL2 and from 84.4% to 91.2% with Qwen2-VL. In LLaVA-OneVision, DocVLM achieves improved results while using 80% less image tokens. The reduced token usage allows processing multiple pages effectively, showing impressive zero-shot results on DUDE and state-of-the-art performance on MP-DocVQA, highlighting DocVLM's potential for applications requiring high-performance and efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08746
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DocVLM: Make Your VLM an Efficient Reader
Nacson, Mor Shpigel
Aberdam, Aviad
Ganz, Roy
Avraham, Elad Ben
Golts, Alona
Kittenplon, Yair
Mazor, Shai
Litman, Ron
Computer Vision and Pattern Recognition
Machine Learning
Vision-Language Models (VLMs) excel in diverse visual tasks but face challenges in document understanding, which requires fine-grained text processing. While typical visual tasks perform well with low-resolution inputs, reading-intensive applications demand high-resolution, resulting in significant computational overhead. Using OCR-extracted text in VLM prompts partially addresses this issue but underperforms compared to full-resolution counterpart, as it lacks the complete visual context needed for optimal performance. We introduce DocVLM, a method that integrates an OCR-based modality into VLMs to enhance document processing while preserving original weights. Our approach employs an OCR encoder to capture textual content and layout, compressing these into a compact set of learned queries incorporated into the VLM. Comprehensive evaluations across leading VLMs show that DocVLM significantly reduces reliance on high-resolution images for document understanding. In limited-token regimes (448$\times$448), DocVLM with 64 learned queries improves DocVQA results from 56.0% to 86.6% when integrated with InternVL2 and from 84.4% to 91.2% with Qwen2-VL. In LLaVA-OneVision, DocVLM achieves improved results while using 80% less image tokens. The reduced token usage allows processing multiple pages effectively, showing impressive zero-shot results on DUDE and state-of-the-art performance on MP-DocVQA, highlighting DocVLM's potential for applications requiring high-performance and efficiency.
title DocVLM: Make Your VLM an Efficient Reader
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2412.08746