From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Santos, José Guilherme Marques dos, Yang, Ricardo, Pereira, Rui Humberto, Sousa, Alexandre, Faria, Brígida Mónica, Cardoso, Henrique Lopes, Duarte, José, Reis, José Luís, Reis, Luís Paulo, Pimenta, Pedro, Santos, José Paulo Marques dos
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911717572739072
author Santos, José Guilherme Marques dos
Yang, Ricardo
Pereira, Rui Humberto
Sousa, Alexandre
Faria, Brígida Mónica
Cardoso, Henrique Lopes
Duarte, José
Reis, José Luís
Reis, Luís Paulo
Pimenta, Pedro
Santos, José Paulo Marques dos
author_facet Santos, José Guilherme Marques dos
Yang, Ricardo
Pereira, Rui Humberto
Sousa, Alexandre
Faria, Brígida Mónica
Cardoso, Henrique Lopes
Duarte, José
Reis, José Luís
Reis, Luís Paulo
Pimenta, Pedro
Santos, José Paulo Marques dos
contents Retrieval-Augmented Generation (RAG) systems depend critically on the quality of document preprocessing, yet no prior study has evaluated PDF processing frameworks by their impact on downstream question-answering accuracy. We address this gap through a systematic comparison of four open-source PDF-to-Markdown conversion frameworks, Docling, MinerU, Marker, and DeepSeek OCR, across 21 pipeline configurations, varying the conversion tool, cleaning transformations, splitting strategy, and metadata enrichment. Evaluation was performed using a 50-question benchmark over a corpus of 36 Portuguese administrative documents (1706 pages, ~492K words), with LLM-as-judge scoring over 50 independent runs per configuration. Statistical significance was assessed via Wilcoxon signed-rank tests with Cohen's d effect sizes. Two baselines bounded the results: naïve PDFLoader (86.2%) and manually curated Markdown (91.3%). Docling with hierarchical splitting and image descriptions achieved the highest automated accuracy (94.1 +/- 1.6%), surpassing even manual curation. A per-question-type analysis revealed that table-dependent questions drive the largest accuracy differences, with a 33-percentage-point gap between basic and hierarchical splitting. Metadata enrichment and hierarchy-aware chunking contributed more to accuracy than the conversion framework alone. An exploratory GraphRAG implementation underperformed basic RAG (82% vs. 94.1%). These findings demonstrate that data preparation quality is the dominant factor in RAG system performance.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04948
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering
Santos, José Guilherme Marques dos
Yang, Ricardo
Pereira, Rui Humberto
Sousa, Alexandre
Faria, Brígida Mónica
Cardoso, Henrique Lopes
Duarte, José
Reis, José Luís
Reis, Luís Paulo
Pimenta, Pedro
Santos, José Paulo Marques dos
Information Retrieval
Artificial Intelligence
Machine Learning
68T50
I.2.7
Retrieval-Augmented Generation (RAG) systems depend critically on the quality of document preprocessing, yet no prior study has evaluated PDF processing frameworks by their impact on downstream question-answering accuracy. We address this gap through a systematic comparison of four open-source PDF-to-Markdown conversion frameworks, Docling, MinerU, Marker, and DeepSeek OCR, across 21 pipeline configurations, varying the conversion tool, cleaning transformations, splitting strategy, and metadata enrichment. Evaluation was performed using a 50-question benchmark over a corpus of 36 Portuguese administrative documents (1706 pages, ~492K words), with LLM-as-judge scoring over 50 independent runs per configuration. Statistical significance was assessed via Wilcoxon signed-rank tests with Cohen's d effect sizes. Two baselines bounded the results: naïve PDFLoader (86.2%) and manually curated Markdown (91.3%). Docling with hierarchical splitting and image descriptions achieved the highest automated accuracy (94.1 +/- 1.6%), surpassing even manual curation. A per-question-type analysis revealed that table-dependent questions drive the largest accuracy differences, with a 33-percentage-point gap between basic and hierarchical splitting. Metadata enrichment and hierarchy-aware chunking contributed more to accuracy than the conversion framework alone. An exploratory GraphRAG implementation underperformed basic RAG (82% vs. 94.1%). These findings demonstrate that data preparation quality is the dominant factor in RAG system performance.
title From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering
topic Information Retrieval
Artificial Intelligence
Machine Learning
68T50
I.2.7
url https://arxiv.org/abs/2604.04948