KVP10k : A Comprehensive Dataset for Key-Value Pair Extraction in Business Documents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Naparstek, Oshri, Pony, Roi, Shapira, Inbar, Dahood, Foad Abo, Azulai, Ophir, Yaroker, Yevgeny, Rubinstein, Nadav, Lysak, Maksym, Staar, Peter, Nassar, Ahmed, Livathinos, Nikolaos, Auer, Christoph, Amrani, Elad, Friedman, Idan, Prince, Orit, Burshtein, Yevgeny, Goldfarb, Adi Raz, Barzelay, Udi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913338253901824
author Naparstek, Oshri
Pony, Roi
Shapira, Inbar
Dahood, Foad Abo
Azulai, Ophir
Yaroker, Yevgeny
Rubinstein, Nadav
Lysak, Maksym
Staar, Peter
Nassar, Ahmed
Livathinos, Nikolaos
Auer, Christoph
Amrani, Elad
Friedman, Idan
Prince, Orit
Burshtein, Yevgeny
Goldfarb, Adi Raz
Barzelay, Udi
author_facet Naparstek, Oshri
Pony, Roi
Shapira, Inbar
Dahood, Foad Abo
Azulai, Ophir
Yaroker, Yevgeny
Rubinstein, Nadav
Lysak, Maksym
Staar, Peter
Nassar, Ahmed
Livathinos, Nikolaos
Auer, Christoph
Amrani, Elad
Friedman, Idan
Prince, Orit
Burshtein, Yevgeny
Goldfarb, Adi Raz
Barzelay, Udi
contents In recent years, the challenge of extracting information from business documents has emerged as a critical task, finding applications across numerous domains. This effort has attracted substantial interest from both industry and academy, highlighting its significance in the current technological landscape. Most datasets in this area are primarily focused on Key Information Extraction (KIE), where the extraction process revolves around extracting information using a specific, predefined set of keys. Unlike most existing datasets and benchmarks, our focus is on discovering key-value pairs (KVPs) without relying on predefined keys, navigating through an array of diverse templates and complex layouts. This task presents unique challenges, primarily due to the absence of comprehensive datasets and benchmarks tailored for non-predetermined KVP extraction. To address this gap, we introduce KVP10k , a new dataset and benchmark specifically designed for KVP extraction. The dataset contains 10707 richly annotated images. In our benchmark, we also introduce a new challenging task that combines elements of KIE as well as KVP in a single task. KVP10k sets itself apart with its extensive diversity in data and richly detailed annotations, paving the way for advancements in the field of information extraction from complex business documents.
format Preprint
id arxiv_https___arxiv_org_abs_2405_00505
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle KVP10k : A Comprehensive Dataset for Key-Value Pair Extraction in Business Documents
Naparstek, Oshri
Pony, Roi
Shapira, Inbar
Dahood, Foad Abo
Azulai, Ophir
Yaroker, Yevgeny
Rubinstein, Nadav
Lysak, Maksym
Staar, Peter
Nassar, Ahmed
Livathinos, Nikolaos
Auer, Christoph
Amrani, Elad
Friedman, Idan
Prince, Orit
Burshtein, Yevgeny
Goldfarb, Adi Raz
Barzelay, Udi
Information Retrieval
Machine Learning
In recent years, the challenge of extracting information from business documents has emerged as a critical task, finding applications across numerous domains. This effort has attracted substantial interest from both industry and academy, highlighting its significance in the current technological landscape. Most datasets in this area are primarily focused on Key Information Extraction (KIE), where the extraction process revolves around extracting information using a specific, predefined set of keys. Unlike most existing datasets and benchmarks, our focus is on discovering key-value pairs (KVPs) without relying on predefined keys, navigating through an array of diverse templates and complex layouts. This task presents unique challenges, primarily due to the absence of comprehensive datasets and benchmarks tailored for non-predetermined KVP extraction. To address this gap, we introduce KVP10k , a new dataset and benchmark specifically designed for KVP extraction. The dataset contains 10707 richly annotated images. In our benchmark, we also introduce a new challenging task that combines elements of KIE as well as KVP in a single task. KVP10k sets itself apart with its extensive diversity in data and richly detailed annotations, paving the way for advancements in the field of information extraction from complex business documents.
title KVP10k : A Comprehensive Dataset for Key-Value Pair Extraction in Business Documents
topic Information Retrieval
Machine Learning
url https://arxiv.org/abs/2405.00505