A Reality Check on Context Utilisation for Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hagström, Lovisa, Marjanović, Sara Vera, Yu, Haeun, Arora, Arnav, Lioma, Christina, Maistro, Maria, Atanasova, Pepa, Augenstein, Isabelle
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916765815013376
author Hagström, Lovisa
Marjanović, Sara Vera
Yu, Haeun
Arora, Arnav
Lioma, Christina
Maistro, Maria
Atanasova, Pepa
Augenstein, Isabelle
author_facet Hagström, Lovisa
Marjanović, Sara Vera
Yu, Haeun
Arora, Arnav
Lioma, Christina
Maistro, Maria
Atanasova, Pepa
Augenstein, Isabelle
contents Retrieval-augmented generation (RAG) helps address the limitations of parametric knowledge embedded within a language model (LM). In real world settings, retrieved information can vary in complexity, yet most investigations of LM utilisation of context has been limited to synthetic text. We introduce DRUID (Dataset of Retrieved Unreliable, Insufficient and Difficult-to-understand contexts) with real-world queries and contexts manually annotated for stance. The dataset is based on the prototypical task of automated claim verification, for which automated retrieval of real-world evidence is crucial. We compare DRUID to synthetic datasets (CounterFact, ConflictQA) and find that artificial datasets often fail to represent the complexity and diversity of realistically retrieved context. We show that synthetic datasets exaggerate context characteristics rare in real retrieved data, which leads to inflated context utilisation results, as measured by our novel ACU score. Moreover, while previous work has mainly focused on singleton context characteristics to explain context utilisation, correlations between singleton context properties and ACU on DRUID are surprisingly small compared to other properties related to context source. Overall, our work underscores the need for real-world aligned context utilisation studies to represent and improve performance in real-world RAG settings.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17031
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Reality Check on Context Utilisation for Retrieval-Augmented Generation
Hagström, Lovisa
Marjanović, Sara Vera
Yu, Haeun
Arora, Arnav
Lioma, Christina
Maistro, Maria
Atanasova, Pepa
Augenstein, Isabelle
Computation and Language
Artificial Intelligence
Retrieval-augmented generation (RAG) helps address the limitations of parametric knowledge embedded within a language model (LM). In real world settings, retrieved information can vary in complexity, yet most investigations of LM utilisation of context has been limited to synthetic text. We introduce DRUID (Dataset of Retrieved Unreliable, Insufficient and Difficult-to-understand contexts) with real-world queries and contexts manually annotated for stance. The dataset is based on the prototypical task of automated claim verification, for which automated retrieval of real-world evidence is crucial. We compare DRUID to synthetic datasets (CounterFact, ConflictQA) and find that artificial datasets often fail to represent the complexity and diversity of realistically retrieved context. We show that synthetic datasets exaggerate context characteristics rare in real retrieved data, which leads to inflated context utilisation results, as measured by our novel ACU score. Moreover, while previous work has mainly focused on singleton context characteristics to explain context utilisation, correlations between singleton context properties and ACU on DRUID are surprisingly small compared to other properties related to context source. Overall, our work underscores the need for real-world aligned context utilisation studies to represent and improve performance in real-world RAG settings.
title A Reality Check on Context Utilisation for Retrieval-Augmented Generation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.17031