Enabling and Analyzing How to Efficiently Extract Information from Hybrid Long Documents with LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yue, Chongjian, Xu, Xinrun, Ma, Xiaojun, Du, Lun, Liu, Hengyu, Ding, Zhiming, Jiang, Yanbing, Han, Shi, Zhang, Dongmei
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917606942834688
author Yue, Chongjian
Xu, Xinrun
Ma, Xiaojun
Du, Lun
Liu, Hengyu
Ding, Zhiming
Jiang, Yanbing
Han, Shi
Zhang, Dongmei
author_facet Yue, Chongjian
Xu, Xinrun
Ma, Xiaojun
Du, Lun
Liu, Hengyu
Ding, Zhiming
Jiang, Yanbing
Han, Shi
Zhang, Dongmei
contents Large Language Models (LLMs) demonstrate exceptional performance in textual understanding and tabular reasoning tasks. However, their ability to comprehend and analyze hybrid text, containing textual and tabular data, remains underexplored. In this research, we specialize in harnessing the potential of LLMs to comprehend critical information from financial reports, which are hybrid long-documents. We propose an Automated Financial Information Extraction (AFIE) framework that enhances LLMs' ability to comprehend and extract information from financial reports. To evaluate AFIE, we develop a Financial Reports Numerical Extraction (FINE) dataset and conduct an extensive experimental analysis. Our framework is effectively validated on GPT-3.5 and GPT-4, yielding average accuracy increases of 53.94% and 33.77%, respectively, compared to a naive method. These results suggest that the AFIE framework offers accuracy for automated numerical extraction from complex, hybrid documents.
format Preprint
id arxiv_https___arxiv_org_abs_2305_16344
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Enabling and Analyzing How to Efficiently Extract Information from Hybrid Long Documents with LLMs
Yue, Chongjian
Xu, Xinrun
Ma, Xiaojun
Du, Lun
Liu, Hengyu
Ding, Zhiming
Jiang, Yanbing
Han, Shi
Zhang, Dongmei
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) demonstrate exceptional performance in textual understanding and tabular reasoning tasks. However, their ability to comprehend and analyze hybrid text, containing textual and tabular data, remains underexplored. In this research, we specialize in harnessing the potential of LLMs to comprehend critical information from financial reports, which are hybrid long-documents. We propose an Automated Financial Information Extraction (AFIE) framework that enhances LLMs' ability to comprehend and extract information from financial reports. To evaluate AFIE, we develop a Financial Reports Numerical Extraction (FINE) dataset and conduct an extensive experimental analysis. Our framework is effectively validated on GPT-3.5 and GPT-4, yielding average accuracy increases of 53.94% and 33.77%, respectively, compared to a naive method. These results suggest that the AFIE framework offers accuracy for automated numerical extraction from complex, hybrid documents.
title Enabling and Analyzing How to Efficiently Extract Information from Hybrid Long Documents with LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2305.16344