Extract Information from Hybrid Long Documents Leveraging LLMs: A Framework and Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yue, Chongjian, Xu, Xinrun, Ma, Xiaojun, Du, Lun, Ding, Zhiming, Han, Shi, Zhang, Dongmei, Zhang, Qi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909445547622400
author Yue, Chongjian
Xu, Xinrun
Ma, Xiaojun
Du, Lun
Ding, Zhiming
Han, Shi
Zhang, Dongmei
Zhang, Qi
author_facet Yue, Chongjian
Xu, Xinrun
Ma, Xiaojun
Du, Lun
Ding, Zhiming
Han, Shi
Zhang, Dongmei
Zhang, Qi
contents Large Language Models (LLMs) demonstrate exceptional performance in textual understanding and tabular reasoning tasks. However, their ability to comprehend and analyze hybrid text, containing textual and tabular data, remains unexplored. The hybrid text often appears in the form of hybrid long documents (HLDs), which far exceed the token limit of LLMs. Consequently, we apply an Automated Information Extraction framework (AIE) to enable LLMs to process the HLDs and carry out experiments to analyse four important aspects of information extraction from HLDs. Given the findings: 1) The effective way to select and summarize the useful part of a HLD. 2) An easy table serialization way is enough for LLMs to understand tables. 3) The naive AIE has adaptability in many complex scenarios. 4) The useful prompt engineering to enhance LLMs on HLDs. To address the issue of dataset scarcity in HLDs and support future work, we also propose the Financial Reports Numerical Extraction (FINE) dataset. The dataset and code are publicly available in the attachments.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20072
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Extract Information from Hybrid Long Documents Leveraging LLMs: A Framework and Dataset
Yue, Chongjian
Xu, Xinrun
Ma, Xiaojun
Du, Lun
Ding, Zhiming
Han, Shi
Zhang, Dongmei
Zhang, Qi
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) demonstrate exceptional performance in textual understanding and tabular reasoning tasks. However, their ability to comprehend and analyze hybrid text, containing textual and tabular data, remains unexplored. The hybrid text often appears in the form of hybrid long documents (HLDs), which far exceed the token limit of LLMs. Consequently, we apply an Automated Information Extraction framework (AIE) to enable LLMs to process the HLDs and carry out experiments to analyse four important aspects of information extraction from HLDs. Given the findings: 1) The effective way to select and summarize the useful part of a HLD. 2) An easy table serialization way is enough for LLMs to understand tables. 3) The naive AIE has adaptability in many complex scenarios. 4) The useful prompt engineering to enhance LLMs on HLDs. To address the issue of dataset scarcity in HLDs and support future work, we also propose the Financial Reports Numerical Extraction (FINE) dataset. The dataset and code are publicly available in the attachments.
title Extract Information from Hybrid Long Documents Leveraging LLMs: A Framework and Dataset
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.20072