RealKIE: Five Novel Datasets for Enterprise Key Information Extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Townsend, Benjamin, May, Madison, Mackowiak, Katherine, Wells, Christopher
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915534074806272
author Townsend, Benjamin
May, Madison
Mackowiak, Katherine
Wells, Christopher
author_facet Townsend, Benjamin
May, Madison
Mackowiak, Katherine
Wells, Christopher
contents We introduce RealKIE, a benchmark of five challenging datasets aimed at advancing key information extraction methods, with an emphasis on enterprise applications. The datasets include a diverse range of documents including SEC S1 Filings, US Non-disclosure Agreements, UK Charity Reports, FCC Invoices, and Resource Contracts. Each presents unique challenges: poor text serialization, sparse annotations in long documents, and complex tabular layouts. These datasets provide a realistic testing ground for key information extraction tasks like investment analysis and contract analysis. In addition to presenting these datasets, we offer an in-depth description of the annotation process, document processing techniques, and baseline modeling approaches. This contribution facilitates the development of NLP models capable of handling practical challenges and supports further research into information extraction technologies applicable to industry-specific problems. The annotated data, OCR outputs, and code to reproduce baselines are available to download at https://indicodatasolutions.github.io/RealKIE/.
format Preprint
id arxiv_https___arxiv_org_abs_2403_20101
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RealKIE: Five Novel Datasets for Enterprise Key Information Extraction
Townsend, Benjamin
May, Madison
Mackowiak, Katherine
Wells, Christopher
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
We introduce RealKIE, a benchmark of five challenging datasets aimed at advancing key information extraction methods, with an emphasis on enterprise applications. The datasets include a diverse range of documents including SEC S1 Filings, US Non-disclosure Agreements, UK Charity Reports, FCC Invoices, and Resource Contracts. Each presents unique challenges: poor text serialization, sparse annotations in long documents, and complex tabular layouts. These datasets provide a realistic testing ground for key information extraction tasks like investment analysis and contract analysis. In addition to presenting these datasets, we offer an in-depth description of the annotation process, document processing techniques, and baseline modeling approaches. This contribution facilitates the development of NLP models capable of handling practical challenges and supports further research into information extraction technologies applicable to industry-specific problems. The annotated data, OCR outputs, and code to reproduce baselines are available to download at https://indicodatasolutions.github.io/RealKIE/.
title RealKIE: Five Novel Datasets for Enterprise Key Information Extraction
topic Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2403.20101