Zero-Shot Prompting and Few-Shot Fine-Tuning: Revisiting Document Image Classification Using Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Scius-Bertrand, Anna, Jungo, Michael, Vögtlin, Lars, Spat, Jean-Marc, Fischer, Andreas
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909433012944896
author Scius-Bertrand, Anna
Jungo, Michael
Vögtlin, Lars
Spat, Jean-Marc
Fischer, Andreas
author_facet Scius-Bertrand, Anna
Jungo, Michael
Vögtlin, Lars
Spat, Jean-Marc
Fischer, Andreas
contents Classifying scanned documents is a challenging problem that involves image, layout, and text analysis for document understanding. Nevertheless, for certain benchmark datasets, notably RVL-CDIP, the state of the art is closing in to near-perfect performance when considering hundreds of thousands of training samples. With the advent of large language models (LLMs), which are excellent few-shot learners, the question arises to what extent the document classification problem can be addressed with only a few training samples, or even none at all. In this paper, we investigate this question in the context of zero-shot prompting and few-shot model fine-tuning, with the aim of reducing the need for human-annotated training samples as much as possible.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13859
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Zero-Shot Prompting and Few-Shot Fine-Tuning: Revisiting Document Image Classification Using Large Language Models
Scius-Bertrand, Anna
Jungo, Michael
Vögtlin, Lars
Spat, Jean-Marc
Fischer, Andreas
Computer Vision and Pattern Recognition
Classifying scanned documents is a challenging problem that involves image, layout, and text analysis for document understanding. Nevertheless, for certain benchmark datasets, notably RVL-CDIP, the state of the art is closing in to near-perfect performance when considering hundreds of thousands of training samples. With the advent of large language models (LLMs), which are excellent few-shot learners, the question arises to what extent the document classification problem can be addressed with only a few training samples, or even none at all. In this paper, we investigate this question in the context of zero-shot prompting and few-shot model fine-tuning, with the aim of reducing the need for human-annotated training samples as much as possible.
title Zero-Shot Prompting and Few-Shot Fine-Tuning: Revisiting Document Image Classification Using Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.13859