Empirical Evaluation of PDF Parsing and Chunking for Financial Question Answering with RAG

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bachyr, Omar El, Song, Yewei, Ezzini, Saad, Klein, Jacques, Bissyandé, Tegawendé F., Zilali, Anas, Ble, Ulrick, Goujon, Anne
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915935979307008
author Bachyr, Omar El
Song, Yewei
Ezzini, Saad
Klein, Jacques
Bissyandé, Tegawendé F.
Zilali, Anas
Ble, Ulrick
Goujon, Anne
author_facet Bachyr, Omar El
Song, Yewei
Ezzini, Saad
Klein, Jacques
Bissyandé, Tegawendé F.
Zilali, Anas
Ble, Ulrick
Goujon, Anne
contents PDF files are primarily intended for human reading rather than automated processing. In addition, the heterogeneous content of PDFs, such as text, tables, and images, poses significant challenges for parsing and information extraction. To address these difficulties, both practitioners and researchers are increasingly developing new methods, including the promising Retrieval-Augmented Generation (RAG) systems to automated PDF processing. However, there is no comprehensive study investigating how different components and design choices affect the performance of a RAG system for understanding PDFs. In this paper, we propose such a study (1) by focusing on Question Answering, a specific language understanding task, and (2) by leveraging two benchmarks from the financial domain, including TableQuest, our newly generated, publicly available benchmark. We systematically examine multiple PDF parsers and chunking strategies (with varied overlap), along with their potential synergies in preserving document structure and ensuring answer correctness. Overall, our results offer practical guidelines for building robust RAG pipelines for PDF understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12047
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Empirical Evaluation of PDF Parsing and Chunking for Financial Question Answering with RAG
Bachyr, Omar El
Song, Yewei
Ezzini, Saad
Klein, Jacques
Bissyandé, Tegawendé F.
Zilali, Anas
Ble, Ulrick
Goujon, Anne
Computation and Language
Information Retrieval
PDF files are primarily intended for human reading rather than automated processing. In addition, the heterogeneous content of PDFs, such as text, tables, and images, poses significant challenges for parsing and information extraction. To address these difficulties, both practitioners and researchers are increasingly developing new methods, including the promising Retrieval-Augmented Generation (RAG) systems to automated PDF processing. However, there is no comprehensive study investigating how different components and design choices affect the performance of a RAG system for understanding PDFs. In this paper, we propose such a study (1) by focusing on Question Answering, a specific language understanding task, and (2) by leveraging two benchmarks from the financial domain, including TableQuest, our newly generated, publicly available benchmark. We systematically examine multiple PDF parsers and chunking strategies (with varied overlap), along with their potential synergies in preserving document structure and ensuring answer correctness. Overall, our results offer practical guidelines for building robust RAG pipelines for PDF understanding.
title Empirical Evaluation of PDF Parsing and Chunking for Financial Question Answering with RAG
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2604.12047