On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Upadhyay, Shivani, Ataey, Messiah, Murtaza, Syed Shariyar, Nie, Yifan, Lin, Jimmy
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911113022537728
author Upadhyay, Shivani
Ataey, Messiah
Murtaza, Syed Shariyar
Nie, Yifan
Lin, Jimmy
author_facet Upadhyay, Shivani
Ataey, Messiah
Murtaza, Syed Shariyar
Nie, Yifan
Lin, Jimmy
contents The proliferation of complex structured data in hybrid sources, such as PDF documents and web pages, presents unique challenges for current Large Language Models (LLMs) and Multi-modal Large Language Models (MLLMs) in providing accurate answers. Despite the recent advancements of MLLMs, they still often falter when interpreting intricately structured information, such as nested tables and multi-dimensional plots, leading to hallucinations and erroneous outputs. This paper explores the capabilities of LLMs and MLLMs in understanding and answering questions from complex data structures found in PDF documents by leveraging industrial and open-source tools as part of a pre-processing pipeline. Our findings indicate that GPT-4o, a popular MLLM, achieves an accuracy of 56% on multi-structured documents when fed documents directly, and that integrating pre-processing tools raises the accuracy of LLMs to 61.3% for GPT-4o and 76% for GPT-4, and with lower overall cost. The code is publicly available at https://github.com/OGCDS/FinancialQA.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05182
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools
Upadhyay, Shivani
Ataey, Messiah
Murtaza, Syed Shariyar
Nie, Yifan
Lin, Jimmy
Information Retrieval
The proliferation of complex structured data in hybrid sources, such as PDF documents and web pages, presents unique challenges for current Large Language Models (LLMs) and Multi-modal Large Language Models (MLLMs) in providing accurate answers. Despite the recent advancements of MLLMs, they still often falter when interpreting intricately structured information, such as nested tables and multi-dimensional plots, leading to hallucinations and erroneous outputs. This paper explores the capabilities of LLMs and MLLMs in understanding and answering questions from complex data structures found in PDF documents by leveraging industrial and open-source tools as part of a pre-processing pipeline. Our findings indicate that GPT-4o, a popular MLLM, achieves an accuracy of 56% on multi-structured documents when fed documents directly, and that integrating pre-processing tools raises the accuracy of LLMs to 61.3% for GPT-4o and 76% for GPT-4, and with lower overall cost. The code is publicly available at https://github.com/OGCDS/FinancialQA.
title On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools
topic Information Retrieval
url https://arxiv.org/abs/2506.05182