To Know or Not To Know? Analyzing Self-Consistency of Large Language Models under Ambiguity

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sedova, Anastasiia, Litschko, Robert, Frassinelli, Diego, Roth, Benjamin, Plank, Barbara
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929527540678656
author Sedova, Anastasiia
Litschko, Robert
Frassinelli, Diego
Roth, Benjamin
Plank, Barbara
author_facet Sedova, Anastasiia
Litschko, Robert
Frassinelli, Diego
Roth, Benjamin
Plank, Barbara
contents One of the major aspects contributing to the striking performance of large language models (LLMs) is the vast amount of factual knowledge accumulated during pre-training. Yet, many LLMs suffer from self-inconsistency, which raises doubts about their trustworthiness and reliability. This paper focuses on entity type ambiguity, analyzing the proficiency and consistency of state-of-the-art LLMs in applying factual knowledge when prompted with ambiguous entities. To do so, we propose an evaluation protocol that disentangles knowing from applying knowledge, and test state-of-the-art LLMs on 49 ambiguous entities. Our experiments reveal that LLMs struggle with choosing the correct entity reading, achieving an average accuracy of only 85%, and as low as 75% with underspecified prompts. The results also reveal systematic discrepancies in LLM behavior, showing that while the models may possess knowledge, they struggle to apply it consistently, exhibit biases toward preferred readings, and display self-inconsistencies. This highlights the need to address entity ambiguity in the future for more trustworthy LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2407_17125
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle To Know or Not To Know? Analyzing Self-Consistency of Large Language Models under Ambiguity
Sedova, Anastasiia
Litschko, Robert
Frassinelli, Diego
Roth, Benjamin
Plank, Barbara
Computation and Language
Machine Learning
One of the major aspects contributing to the striking performance of large language models (LLMs) is the vast amount of factual knowledge accumulated during pre-training. Yet, many LLMs suffer from self-inconsistency, which raises doubts about their trustworthiness and reliability. This paper focuses on entity type ambiguity, analyzing the proficiency and consistency of state-of-the-art LLMs in applying factual knowledge when prompted with ambiguous entities. To do so, we propose an evaluation protocol that disentangles knowing from applying knowledge, and test state-of-the-art LLMs on 49 ambiguous entities. Our experiments reveal that LLMs struggle with choosing the correct entity reading, achieving an average accuracy of only 85%, and as low as 75% with underspecified prompts. The results also reveal systematic discrepancies in LLM behavior, showing that while the models may possess knowledge, they struggle to apply it consistently, exhibit biases toward preferred readings, and display self-inconsistencies. This highlights the need to address entity ambiguity in the future for more trustworthy LLMs.
title To Know or Not To Know? Analyzing Self-Consistency of Large Language Models under Ambiguity
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2407.17125