TabRAG: Improving Tabular Document Question Answering for Retrieval Augmented Generation via Structured Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Si, Jacob, Qu, Mike, Lee, Michelle, Rei, Marek, Li, Yingzhen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915765839462400
author Si, Jacob
Qu, Mike
Lee, Michelle
Rei, Marek
Li, Yingzhen
author_facet Si, Jacob
Qu, Mike
Lee, Michelle
Rei, Marek
Li, Yingzhen
contents Incorporating external knowledge bases in traditional retrieval-augmented generation (RAG) relies on parsing the document, followed by querying a language model with the parsed information via in-context learning. While effective for text-based documents, question answering on tabular documents often fails to generate plausible responses. Standard parsing techniques lose the two-dimensional structural semantics critical for cell interpretation. In this work, we present TabRAG, a parsing-based RAG framework designed to improve tabular document question answering via structured representations. Our framework consists of layout segmentation that decomposes the document inputs into a series of components, enabling fine-grained extraction. Subsequently, a vision language model parses and extracts the document tables into a hierarchically structured representation. In order to cater various table styles and formats, we integrate a self-generated in-context learning module that guides the table extraction process. Experimental results demonstrate that TabRAG outperforms existing popular parsing techniques across a broad suite of evaluation and ablation benchmarks. Code is available at: https://github.com/jacobyhsi/TabRAG.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06582
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TabRAG: Improving Tabular Document Question Answering for Retrieval Augmented Generation via Structured Representations
Si, Jacob
Qu, Mike
Lee, Michelle
Rei, Marek
Li, Yingzhen
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
Machine Learning
Incorporating external knowledge bases in traditional retrieval-augmented generation (RAG) relies on parsing the document, followed by querying a language model with the parsed information via in-context learning. While effective for text-based documents, question answering on tabular documents often fails to generate plausible responses. Standard parsing techniques lose the two-dimensional structural semantics critical for cell interpretation. In this work, we present TabRAG, a parsing-based RAG framework designed to improve tabular document question answering via structured representations. Our framework consists of layout segmentation that decomposes the document inputs into a series of components, enabling fine-grained extraction. Subsequently, a vision language model parses and extracts the document tables into a hierarchically structured representation. In order to cater various table styles and formats, we integrate a self-generated in-context learning module that guides the table extraction process. Experimental results demonstrate that TabRAG outperforms existing popular parsing techniques across a broad suite of evaluation and ablation benchmarks. Code is available at: https://github.com/jacobyhsi/TabRAG.
title TabRAG: Improving Tabular Document Question Answering for Retrieval Augmented Generation via Structured Representations
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2511.06582