Assessing "Implicit" Retrieval Robustness of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Xiaoyu, Blloshmi, Rexhina, Zhu, Dawei, Pei, Jiahuan, Zhang, Wei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909231727247360
author Shen, Xiaoyu
Blloshmi, Rexhina
Zhu, Dawei
Pei, Jiahuan
Zhang, Wei
author_facet Shen, Xiaoyu
Blloshmi, Rexhina
Zhu, Dawei
Pei, Jiahuan
Zhang, Wei
contents Retrieval-augmented generation has gained popularity as a framework to enhance large language models with external knowledge. However, its effectiveness hinges on the retrieval robustness of the model. If the model lacks retrieval robustness, its performance is constrained by the accuracy of the retriever, resulting in significant compromises when the retrieved context is irrelevant. In this paper, we evaluate the "implicit" retrieval robustness of various large language models, instructing them to directly output the final answer without explicitly judging the relevance of the retrieved context. Our findings reveal that fine-tuning on a mix of gold and distracting context significantly enhances the model's robustness to retrieval inaccuracies, while still maintaining its ability to extract correct answers when retrieval is accurate. This suggests that large language models can implicitly handle relevant or irrelevant retrieved context by learning solely from the supervision of the final answer in an end-to-end manner. Introducing an additional process for explicit relevance judgment can be unnecessary and disrupts the end-to-end approach.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18134
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Assessing "Implicit" Retrieval Robustness of Large Language Models
Shen, Xiaoyu
Blloshmi, Rexhina
Zhu, Dawei
Pei, Jiahuan
Zhang, Wei
Computation and Language
Retrieval-augmented generation has gained popularity as a framework to enhance large language models with external knowledge. However, its effectiveness hinges on the retrieval robustness of the model. If the model lacks retrieval robustness, its performance is constrained by the accuracy of the retriever, resulting in significant compromises when the retrieved context is irrelevant. In this paper, we evaluate the "implicit" retrieval robustness of various large language models, instructing them to directly output the final answer without explicitly judging the relevance of the retrieved context. Our findings reveal that fine-tuning on a mix of gold and distracting context significantly enhances the model's robustness to retrieval inaccuracies, while still maintaining its ability to extract correct answers when retrieval is accurate. This suggests that large language models can implicitly handle relevant or irrelevant retrieved context by learning solely from the supervision of the final answer in an end-to-end manner. Introducing an additional process for explicit relevance judgment can be unnecessary and disrupts the end-to-end approach.
title Assessing "Implicit" Retrieval Robustness of Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2406.18134