Say Less, Mean More: Leveraging Pragmatics in Retrieval-Augmented Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Riaz, Haris, Riloff, Ellen, Surdeanu, Mihai
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913709738164224
author Riaz, Haris
Riloff, Ellen
Surdeanu, Mihai
author_facet Riaz, Haris
Riloff, Ellen
Surdeanu, Mihai
contents We propose a simple, unsupervised method that injects pragmatic principles in retrieval-augmented generation (RAG) frameworks such as Dense Passage Retrieval to enhance the utility of retrieved contexts. Our approach first identifies which sentences in a pool of documents retrieved by RAG are most relevant to the question at hand, cover all the topics addressed in the input question and no more, and then highlights these sentences within their context, before they are provided to the LLM, without truncating or altering the context in any other way. We show that this simple idea brings consistent improvements in experiments on three question answering tasks (ARC-Challenge, PubHealth and PopQA) using five different LLMs. It notably enhances relative accuracy by up to 19.7% on PubHealth and 10% on ARC-Challenge compared to a conventional RAG system.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17839
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Say Less, Mean More: Leveraging Pragmatics in Retrieval-Augmented Generation
Riaz, Haris
Riloff, Ellen
Surdeanu, Mihai
Computation and Language
Artificial Intelligence
Machine Learning
We propose a simple, unsupervised method that injects pragmatic principles in retrieval-augmented generation (RAG) frameworks such as Dense Passage Retrieval to enhance the utility of retrieved contexts. Our approach first identifies which sentences in a pool of documents retrieved by RAG are most relevant to the question at hand, cover all the topics addressed in the input question and no more, and then highlights these sentences within their context, before they are provided to the LLM, without truncating or altering the context in any other way. We show that this simple idea brings consistent improvements in experiments on three question answering tasks (ARC-Challenge, PubHealth and PopQA) using five different LLMs. It notably enhances relative accuracy by up to 19.7% on PubHealth and 10% on ARC-Challenge compared to a conventional RAG system.
title Say Less, Mean More: Leveraging Pragmatics in Retrieval-Augmented Generation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2502.17839