From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911717572739072 |
|---|---|
| author | Santos, José Guilherme Marques dos Yang, Ricardo Pereira, Rui Humberto Sousa, Alexandre Faria, Brígida Mónica Cardoso, Henrique Lopes Duarte, José Reis, José Luís Reis, Luís Paulo Pimenta, Pedro Santos, José Paulo Marques dos |
| author_facet | Santos, José Guilherme Marques dos Yang, Ricardo Pereira, Rui Humberto Sousa, Alexandre Faria, Brígida Mónica Cardoso, Henrique Lopes Duarte, José Reis, José Luís Reis, Luís Paulo Pimenta, Pedro Santos, José Paulo Marques dos |
| contents | Retrieval-Augmented Generation (RAG) systems depend critically on the quality of document preprocessing, yet no prior study has evaluated PDF processing frameworks by their impact on downstream question-answering accuracy. We address this gap through a systematic comparison of four open-source PDF-to-Markdown conversion frameworks, Docling, MinerU, Marker, and DeepSeek OCR, across 21 pipeline configurations, varying the conversion tool, cleaning transformations, splitting strategy, and metadata enrichment. Evaluation was performed using a 50-question benchmark over a corpus of 36 Portuguese administrative documents (1706 pages, ~492K words), with LLM-as-judge scoring over 50 independent runs per configuration. Statistical significance was assessed via Wilcoxon signed-rank tests with Cohen's d effect sizes. Two baselines bounded the results: naïve PDFLoader (86.2%) and manually curated Markdown (91.3%). Docling with hierarchical splitting and image descriptions achieved the highest automated accuracy (94.1 +/- 1.6%), surpassing even manual curation. A per-question-type analysis revealed that table-dependent questions drive the largest accuracy differences, with a 33-percentage-point gap between basic and hierarchical splitting. Metadata enrichment and hierarchy-aware chunking contributed more to accuracy than the conversion framework alone. An exploratory GraphRAG implementation underperformed basic RAG (82% vs. 94.1%). These findings demonstrate that data preparation quality is the dominant factor in RAG system performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_04948 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering Santos, José Guilherme Marques dos Yang, Ricardo Pereira, Rui Humberto Sousa, Alexandre Faria, Brígida Mónica Cardoso, Henrique Lopes Duarte, José Reis, José Luís Reis, Luís Paulo Pimenta, Pedro Santos, José Paulo Marques dos Information Retrieval Artificial Intelligence Machine Learning 68T50 I.2.7 Retrieval-Augmented Generation (RAG) systems depend critically on the quality of document preprocessing, yet no prior study has evaluated PDF processing frameworks by their impact on downstream question-answering accuracy. We address this gap through a systematic comparison of four open-source PDF-to-Markdown conversion frameworks, Docling, MinerU, Marker, and DeepSeek OCR, across 21 pipeline configurations, varying the conversion tool, cleaning transformations, splitting strategy, and metadata enrichment. Evaluation was performed using a 50-question benchmark over a corpus of 36 Portuguese administrative documents (1706 pages, ~492K words), with LLM-as-judge scoring over 50 independent runs per configuration. Statistical significance was assessed via Wilcoxon signed-rank tests with Cohen's d effect sizes. Two baselines bounded the results: naïve PDFLoader (86.2%) and manually curated Markdown (91.3%). Docling with hierarchical splitting and image descriptions achieved the highest automated accuracy (94.1 +/- 1.6%), surpassing even manual curation. A per-question-type analysis revealed that table-dependent questions drive the largest accuracy differences, with a 33-percentage-point gap between basic and hierarchical splitting. Metadata enrichment and hierarchy-aware chunking contributed more to accuracy than the conversion framework alone. An exploratory GraphRAG implementation underperformed basic RAG (82% vs. 94.1%). These findings demonstrate that data preparation quality is the dominant factor in RAG system performance. |
| title | From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering |
| topic | Information Retrieval Artificial Intelligence Machine Learning 68T50 I.2.7 |
| url | https://arxiv.org/abs/2604.04948 |