Provence: efficient and robust context pruning for retrieval-augmented generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chirkova, Nadezhda, Formal, Thibault, Nikoulina, Vassilina, Clinchant, Stéphane
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916590038024192
author Chirkova, Nadezhda
Formal, Thibault
Nikoulina, Vassilina
Clinchant, Stéphane
author_facet Chirkova, Nadezhda
Formal, Thibault
Nikoulina, Vassilina
Clinchant, Stéphane
contents Retrieval-augmented generation improves various aspects of large language models (LLMs) generation, but suffers from computational overhead caused by long contexts as well as the propagation of irrelevant retrieved information into generated responses. Context pruning deals with both aspects, by removing irrelevant parts of retrieved contexts before LLM generation. Existing context pruning approaches are however limited, and do not provide a universal model that would be both efficient and robust in a wide range of scenarios, e.g., when contexts contain a variable amount of relevant information or vary in length, or when evaluated on various domains. In this work, we close this gap and introduce Provence (Pruning and Reranking Of retrieVEd relevaNt ContExts), an efficient and robust context pruner for Question Answering, which dynamically detects the needed amount of pruning for a given context and can be used out-of-the-box for various domains. The three key ingredients of Provence are formulating the context pruning task as sequence labeling, unifying context pruning capabilities with context reranking, and training on diverse data. Our experimental results show that Provence enables context pruning with negligible to no drop in performance, in various domains and settings, at almost no cost in a standard RAG pipeline. We also conduct a deeper analysis alongside various ablations to provide insights into training context pruners for future work.
format Preprint
id arxiv_https___arxiv_org_abs_2501_16214
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Provence: efficient and robust context pruning for retrieval-augmented generation
Chirkova, Nadezhda
Formal, Thibault
Nikoulina, Vassilina
Clinchant, Stéphane
Computation and Language
Information Retrieval
Retrieval-augmented generation improves various aspects of large language models (LLMs) generation, but suffers from computational overhead caused by long contexts as well as the propagation of irrelevant retrieved information into generated responses. Context pruning deals with both aspects, by removing irrelevant parts of retrieved contexts before LLM generation. Existing context pruning approaches are however limited, and do not provide a universal model that would be both efficient and robust in a wide range of scenarios, e.g., when contexts contain a variable amount of relevant information or vary in length, or when evaluated on various domains. In this work, we close this gap and introduce Provence (Pruning and Reranking Of retrieVEd relevaNt ContExts), an efficient and robust context pruner for Question Answering, which dynamically detects the needed amount of pruning for a given context and can be used out-of-the-box for various domains. The three key ingredients of Provence are formulating the context pruning task as sequence labeling, unifying context pruning capabilities with context reranking, and training on diverse data. Our experimental results show that Provence enables context pruning with negligible to no drop in performance, in various domains and settings, at almost no cost in a standard RAG pipeline. We also conduct a deeper analysis alongside various ablations to provide insights into training context pruners for future work.
title Provence: efficient and robust context pruning for retrieval-augmented generation
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2501.16214