Notes on Applicability of GPT-4 to Document Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Borchmann, Łukasz
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910461138567168
author Borchmann, Łukasz
author_facet Borchmann, Łukasz
contents We perform a missing, reproducible evaluation of all publicly available GPT-4 family models concerning the Document Understanding field, where it is frequently required to comprehend text spacial arrangement and visual clues in addition to textual semantics. Benchmark results indicate that though it is hard to achieve satisfactory results with text-only models, GPT-4 Vision Turbo performs well when one provides both text recognized by an external OCR engine and document images on the input. Evaluation is followed by analyses that suggest possible contamination of textual GPT-4 models and indicate the significant performance drop for lengthy documents.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18433
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Notes on Applicability of GPT-4 to Document Understanding
Borchmann, Łukasz
Computation and Language
We perform a missing, reproducible evaluation of all publicly available GPT-4 family models concerning the Document Understanding field, where it is frequently required to comprehend text spacial arrangement and visual clues in addition to textual semantics. Benchmark results indicate that though it is hard to achieve satisfactory results with text-only models, GPT-4 Vision Turbo performs well when one provides both text recognized by an external OCR engine and document images on the input. Evaluation is followed by analyses that suggest possible contamination of textual GPT-4 models and indicate the significant performance drop for lengthy documents.
title Notes on Applicability of GPT-4 to Document Understanding
topic Computation and Language
url https://arxiv.org/abs/2405.18433