A Reality Check on Context Utilisation for Retrieval-Augmented Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916765815013376 |
|---|---|
| author | Hagström, Lovisa Marjanović, Sara Vera Yu, Haeun Arora, Arnav Lioma, Christina Maistro, Maria Atanasova, Pepa Augenstein, Isabelle |
| author_facet | Hagström, Lovisa Marjanović, Sara Vera Yu, Haeun Arora, Arnav Lioma, Christina Maistro, Maria Atanasova, Pepa Augenstein, Isabelle |
| contents | Retrieval-augmented generation (RAG) helps address the limitations of parametric knowledge embedded within a language model (LM). In real world settings, retrieved information can vary in complexity, yet most investigations of LM utilisation of context has been limited to synthetic text. We introduce DRUID (Dataset of Retrieved Unreliable, Insufficient and Difficult-to-understand contexts) with real-world queries and contexts manually annotated for stance. The dataset is based on the prototypical task of automated claim verification, for which automated retrieval of real-world evidence is crucial. We compare DRUID to synthetic datasets (CounterFact, ConflictQA) and find that artificial datasets often fail to represent the complexity and diversity of realistically retrieved context. We show that synthetic datasets exaggerate context characteristics rare in real retrieved data, which leads to inflated context utilisation results, as measured by our novel ACU score. Moreover, while previous work has mainly focused on singleton context characteristics to explain context utilisation, correlations between singleton context properties and ACU on DRUID are surprisingly small compared to other properties related to context source. Overall, our work underscores the need for real-world aligned context utilisation studies to represent and improve performance in real-world RAG settings. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_17031 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A Reality Check on Context Utilisation for Retrieval-Augmented Generation Hagström, Lovisa Marjanović, Sara Vera Yu, Haeun Arora, Arnav Lioma, Christina Maistro, Maria Atanasova, Pepa Augenstein, Isabelle Computation and Language Artificial Intelligence Retrieval-augmented generation (RAG) helps address the limitations of parametric knowledge embedded within a language model (LM). In real world settings, retrieved information can vary in complexity, yet most investigations of LM utilisation of context has been limited to synthetic text. We introduce DRUID (Dataset of Retrieved Unreliable, Insufficient and Difficult-to-understand contexts) with real-world queries and contexts manually annotated for stance. The dataset is based on the prototypical task of automated claim verification, for which automated retrieval of real-world evidence is crucial. We compare DRUID to synthetic datasets (CounterFact, ConflictQA) and find that artificial datasets often fail to represent the complexity and diversity of realistically retrieved context. We show that synthetic datasets exaggerate context characteristics rare in real retrieved data, which leads to inflated context utilisation results, as measured by our novel ACU score. Moreover, while previous work has mainly focused on singleton context characteristics to explain context utilisation, correlations between singleton context properties and ACU on DRUID are surprisingly small compared to other properties related to context source. Overall, our work underscores the need for real-world aligned context utilisation studies to represent and improve performance in real-world RAG settings. |
| title | A Reality Check on Context Utilisation for Retrieval-Augmented Generation |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2412.17031 |