Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mündler, Niels, He, Jingxuan, Jenko, Slobodan, Vechev, Martin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913266897256448
author Mündler, Niels
He, Jingxuan
Jenko, Slobodan
Vechev, Martin
author_facet Mündler, Niels
He, Jingxuan
Jenko, Slobodan
Vechev, Martin
contents Large language models (large LMs) are susceptible to producing text that contains hallucinated content. An important instance of this problem is self-contradiction, where the LM generates two contradictory sentences within the same context. In this work, we present a comprehensive investigation into self-contradiction for various instruction-tuned LMs, covering evaluation, detection, and mitigation. Our primary evaluation task is open-domain text generation, but we also demonstrate the applicability of our approach to shorter question answering. Our analysis reveals the prevalence of self-contradictions, e.g., in 17.7% of all sentences produced by ChatGPT. We then propose a novel prompting-based framework designed to effectively detect and mitigate self-contradictions. Our detector achieves high accuracy, e.g., around 80% F1 score when prompting ChatGPT. The mitigation algorithm iteratively refines the generated text to remove contradictory information while preserving text fluency and informativeness. Importantly, our entire framework is applicable to black-box LMs and does not require retrieval of external knowledge. Rather, our method complements retrieval-based methods, as a large portion of self-contradictions (e.g., 35.2% for ChatGPT) cannot be verified using online text. Our approach is practically effective and has been released as a push-button tool to benefit the public at https://chatprotect.ai/.
format Preprint
id arxiv_https___arxiv_org_abs_2305_15852
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
Mündler, Niels
He, Jingxuan
Jenko, Slobodan
Vechev, Martin
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (large LMs) are susceptible to producing text that contains hallucinated content. An important instance of this problem is self-contradiction, where the LM generates two contradictory sentences within the same context. In this work, we present a comprehensive investigation into self-contradiction for various instruction-tuned LMs, covering evaluation, detection, and mitigation. Our primary evaluation task is open-domain text generation, but we also demonstrate the applicability of our approach to shorter question answering. Our analysis reveals the prevalence of self-contradictions, e.g., in 17.7% of all sentences produced by ChatGPT. We then propose a novel prompting-based framework designed to effectively detect and mitigate self-contradictions. Our detector achieves high accuracy, e.g., around 80% F1 score when prompting ChatGPT. The mitigation algorithm iteratively refines the generated text to remove contradictory information while preserving text fluency and informativeness. Importantly, our entire framework is applicable to black-box LMs and does not require retrieval of external knowledge. Rather, our method complements retrieval-based methods, as a large portion of self-contradictions (e.g., 35.2% for ChatGPT) cannot be verified using online text. Our approach is practically effective and has been released as a push-button tool to benefit the public at https://chatprotect.ai/.
title Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2305.15852