TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fhima, Jonathan, Avraham, Elad Ben, Nuriel, Oren, Kittenplon, Yair, Ganz, Roy, Aberdam, Aviad, Litman, Ron
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909379724312576
author Fhima, Jonathan
Avraham, Elad Ben
Nuriel, Oren
Kittenplon, Yair
Ganz, Roy
Aberdam, Aviad
Litman, Ron
author_facet Fhima, Jonathan
Avraham, Elad Ben
Nuriel, Oren
Kittenplon, Yair
Ganz, Roy
Aberdam, Aviad
Litman, Ron
contents Vision-Language (VL) models have garnered considerable research interest; however, they still face challenges in effectively handling text within images. To address this limitation, researchers have developed two approaches. The first method involves utilizing external Optical Character Recognition (OCR) tools to extract textual information from images, which is then prepended to other textual inputs. The second strategy focuses on employing extremely high-resolution images to improve text recognition capabilities. In this paper, we focus on enhancing the first strategy by introducing a novel method, named TAP-VL, which treats OCR information as a distinct modality and seamlessly integrates it into any VL model. TAP-VL employs a lightweight transformer-based OCR module to receive OCR with layout information, compressing it into a short fixed-length sequence for input into the LLM. Initially, we conduct model-agnostic pretraining of the OCR module on unlabeled documents, followed by its integration into any VL architecture through brief fine-tuning. Extensive experiments demonstrate consistent performance improvements when applying TAP-VL to top-performing VL models, across scene-text and document-based VL benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2411_04642
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models
Fhima, Jonathan
Avraham, Elad Ben
Nuriel, Oren
Kittenplon, Yair
Ganz, Roy
Aberdam, Aviad
Litman, Ron
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision-Language (VL) models have garnered considerable research interest; however, they still face challenges in effectively handling text within images. To address this limitation, researchers have developed two approaches. The first method involves utilizing external Optical Character Recognition (OCR) tools to extract textual information from images, which is then prepended to other textual inputs. The second strategy focuses on employing extremely high-resolution images to improve text recognition capabilities. In this paper, we focus on enhancing the first strategy by introducing a novel method, named TAP-VL, which treats OCR information as a distinct modality and seamlessly integrates it into any VL model. TAP-VL employs a lightweight transformer-based OCR module to receive OCR with layout information, compressing it into a short fixed-length sequence for input into the LLM. Initially, we conduct model-agnostic pretraining of the OCR module on unlabeled documents, followed by its integration into any VL architecture through brief fine-tuning. Extensive experiments demonstrate consistent performance improvements when applying TAP-VL to top-performing VL models, across scene-text and document-based VL benchmarks.
title TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.04642