Passage Segmentation of Documents for Extractive Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zuhong, Simon, Charles-Elie, Caspani, Fabien
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913654612426752
author Liu, Zuhong
Simon, Charles-Elie
Caspani, Fabien
author_facet Liu, Zuhong
Simon, Charles-Elie
Caspani, Fabien
contents Retrieval-Augmented Generation (RAG) has proven effective in open-domain question answering. However, the chunking process, which is essential to this pipeline, often receives insufficient attention relative to retrieval and synthesis components. This study emphasizes the critical role of chunking in improving the performance of both dense passage retrieval and the end-to-end RAG pipeline. We then introduce the Logits-Guided Multi-Granular Chunker (LGMGC), a novel framework that splits long documents into contextualized, self-contained chunks of varied granularity. Our experimental results, evaluated on two benchmark datasets, demonstrate that LGMGC not only improves the retrieval step but also outperforms existing chunking methods when integrated into a RAG pipeline.
format Preprint
id arxiv_https___arxiv_org_abs_2501_09940
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Passage Segmentation of Documents for Extractive Question Answering
Liu, Zuhong
Simon, Charles-Elie
Caspani, Fabien
Computation and Language
Information Retrieval
Retrieval-Augmented Generation (RAG) has proven effective in open-domain question answering. However, the chunking process, which is essential to this pipeline, often receives insufficient attention relative to retrieval and synthesis components. This study emphasizes the critical role of chunking in improving the performance of both dense passage retrieval and the end-to-end RAG pipeline. We then introduce the Logits-Guided Multi-Granular Chunker (LGMGC), a novel framework that splits long documents into contextualized, self-contained chunks of varied granularity. Our experimental results, evaluated on two benchmark datasets, demonstrate that LGMGC not only improves the retrieval step but also outperforms existing chunking methods when integrated into a RAG pipeline.
title Passage Segmentation of Documents for Extractive Question Answering
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2501.09940