"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zmigrod, Ran, Shetty, Pranav, Sibue, Mathieu, Ma, Zhiqiang, Nourbakhsh, Armineh, Liu, Xiaomo, Veloso, Manuela
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916497830445056
author Zmigrod, Ran
Shetty, Pranav
Sibue, Mathieu
Ma, Zhiqiang
Nourbakhsh, Armineh
Liu, Xiaomo
Veloso, Manuela
author_facet Zmigrod, Ran
Shetty, Pranav
Sibue, Mathieu
Ma, Zhiqiang
Nourbakhsh, Armineh
Liu, Xiaomo
Veloso, Manuela
contents The rise of large language models (LLMs) for visually rich document understanding (VRDU) has kindled a need for prompt-response, document-based datasets. As annotating new datasets from scratch is labor-intensive, the existing literature has generated prompt-response datasets from available resources using simple templates. For the case of key information extraction (KIE), one of the most common VRDU tasks, past work has typically employed the template "What is the value for the {key}?". However, given the variety of questions encountered in the wild, simple and uniform templates are insufficient for creating robust models in research and industrial contexts. In this work, we present K2Q, a diverse collection of five datasets converted from KIE to a prompt-response format using a plethora of bespoke templates. The questions in K2Q can span multiple entities and be extractive or boolean. We empirically compare the performance of seven baseline generative models on K2Q with zero-shot prompting. We further compare three of these models when training on K2Q versus training on simpler templates to motivate the need of our work. We find that creating diverse and intricate KIE questions enhances the performance and robustness of VRDU models. We hope this work encourages future studies on data quality for generative model training.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15484
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle "What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
Zmigrod, Ran
Shetty, Pranav
Sibue, Mathieu
Ma, Zhiqiang
Nourbakhsh, Armineh
Liu, Xiaomo
Veloso, Manuela
Computation and Language
The rise of large language models (LLMs) for visually rich document understanding (VRDU) has kindled a need for prompt-response, document-based datasets. As annotating new datasets from scratch is labor-intensive, the existing literature has generated prompt-response datasets from available resources using simple templates. For the case of key information extraction (KIE), one of the most common VRDU tasks, past work has typically employed the template "What is the value for the {key}?". However, given the variety of questions encountered in the wild, simple and uniform templates are insufficient for creating robust models in research and industrial contexts. In this work, we present K2Q, a diverse collection of five datasets converted from KIE to a prompt-response format using a plethora of bespoke templates. The questions in K2Q can span multiple entities and be extractive or boolean. We empirically compare the performance of seven baseline generative models on K2Q with zero-shot prompting. We further compare three of these models when training on K2Q versus training on simpler templates to motivate the need of our work. We find that creating diverse and intricate KIE questions enhances the performance and robustness of VRDU models. We hope this work encourages future studies on data quality for generative model training.
title "What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
topic Computation and Language
url https://arxiv.org/abs/2410.15484