PDF Retrieval Augmented Question Answering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hoang, Thi Thu Uyen, Rajendran, Meenakshi, Zhang, Kun, Wu, Yuhan, Nguyen, Viet Anh
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911570735398912
author Hoang, Thi Thu Uyen
Rajendran, Meenakshi
Zhang, Kun
Wu, Yuhan
Nguyen, Viet Anh
author_facet Hoang, Thi Thu Uyen
Rajendran, Meenakshi
Zhang, Kun
Wu, Yuhan
Nguyen, Viet Anh
contents This paper presents an advancement in Question-Answering (QA) systems using a Retrieval Augmented Generation (RAG) framework to enhance information extraction from PDF files. Recognizing the richness and diversity of data within PDFs--including text, images, vector diagrams, graphs, and tables--poses unique challenges for existing QA systems primarily designed for textual content. We seek to develop a comprehensive RAG-based QA system that will effectively address complex multimodal questions, where several data types are combined in the query. This is mainly achieved by refining approaches to processing and integrating non-textual elements in PDFs into the RAG framework to derive precise and relevant answers, as well as fine-tuning large language models to better adapt to our system. We provide an in-depth experimental evaluation of our solution, demonstrating its capability to extract accurate information that can be applied to different types of content across PDFs. This work not only pushes the boundaries of retrieval-augmented QA systems but also lays a foundation for further research in multimodal data integration and processing.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18027
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PDF Retrieval Augmented Question Answering
Hoang, Thi Thu Uyen
Rajendran, Meenakshi
Zhang, Kun
Wu, Yuhan
Nguyen, Viet Anh
Computation and Language
This paper presents an advancement in Question-Answering (QA) systems using a Retrieval Augmented Generation (RAG) framework to enhance information extraction from PDF files. Recognizing the richness and diversity of data within PDFs--including text, images, vector diagrams, graphs, and tables--poses unique challenges for existing QA systems primarily designed for textual content. We seek to develop a comprehensive RAG-based QA system that will effectively address complex multimodal questions, where several data types are combined in the query. This is mainly achieved by refining approaches to processing and integrating non-textual elements in PDFs into the RAG framework to derive precise and relevant answers, as well as fine-tuning large language models to better adapt to our system. We provide an in-depth experimental evaluation of our solution, demonstrating its capability to extract accurate information that can be applied to different types of content across PDFs. This work not only pushes the boundaries of retrieval-augmented QA systems but also lays a foundation for further research in multimodal data integration and processing.
title PDF Retrieval Augmented Question Answering
topic Computation and Language
url https://arxiv.org/abs/2506.18027