Who's Who: Large Language Models Meet Knowledge Conflicts in Practice

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pham, Quang Hieu, Ngo, Hoang, Luu, Anh Tuan, Nguyen, Dat Quoc
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912412184084480
author Pham, Quang Hieu
Ngo, Hoang
Luu, Anh Tuan
Nguyen, Dat Quoc
author_facet Pham, Quang Hieu
Ngo, Hoang
Luu, Anh Tuan
Nguyen, Dat Quoc
contents Retrieval-augmented generation (RAG) methods are viable solutions for addressing the static memory limits of pre-trained language models. Nevertheless, encountering conflicting sources of information within the retrieval context is an inevitable practical challenge. In such situations, the language models are recommended to transparently inform users about the conflicts rather than autonomously deciding what to present based on their inherent biases. To analyze how current large language models (LLMs) align with our recommendation, we introduce WhoQA, a public benchmark dataset to examine model's behavior in knowledge conflict situations. We induce conflicts by asking about a common property among entities having the same name, resulting in questions with up to 8 distinctive answers. WhoQA evaluation set includes 5K questions across 13 Wikidata property types and 150K Wikipedia entities. Our experiments show that despite the simplicity of WhoQA questions, knowledge conflicts significantly degrades LLMs' performance in RAG settings.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15737
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Who's Who: Large Language Models Meet Knowledge Conflicts in Practice
Pham, Quang Hieu
Ngo, Hoang
Luu, Anh Tuan
Nguyen, Dat Quoc
Computation and Language
Artificial Intelligence
Information Retrieval
Retrieval-augmented generation (RAG) methods are viable solutions for addressing the static memory limits of pre-trained language models. Nevertheless, encountering conflicting sources of information within the retrieval context is an inevitable practical challenge. In such situations, the language models are recommended to transparently inform users about the conflicts rather than autonomously deciding what to present based on their inherent biases. To analyze how current large language models (LLMs) align with our recommendation, we introduce WhoQA, a public benchmark dataset to examine model's behavior in knowledge conflict situations. We induce conflicts by asking about a common property among entities having the same name, resulting in questions with up to 8 distinctive answers. WhoQA evaluation set includes 5K questions across 13 Wikidata property types and 150K Wikipedia entities. Our experiments show that despite the simplicity of WhoQA questions, knowledge conflicts significantly degrades LLMs' performance in RAG settings.
title Who's Who: Large Language Models Meet Knowledge Conflicts in Practice
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2410.15737