RELIC: Investigating Large Language Model Responses using Self-Consistency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Furui, Zouhar, Vilém, Arora, Simran, Sachan, Mrinmaya, Strobelt, Hendrik, El-Assady, Mennatallah
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911827826311168
author Cheng, Furui
Zouhar, Vilém
Arora, Simran
Sachan, Mrinmaya
Strobelt, Hendrik
El-Assady, Mennatallah
author_facet Cheng, Furui
Zouhar, Vilém
Arora, Simran
Sachan, Mrinmaya
Strobelt, Hendrik
El-Assady, Mennatallah
contents Large Language Models (LLMs) are notorious for blending fact with fiction and generating non-factual content, known as hallucinations. To address this challenge, we propose an interactive system that helps users gain insight into the reliability of the generated text. Our approach is based on the idea that the self-consistency of multiple samples generated by the same LLM relates to its confidence in individual claims in the generated texts. Using this idea, we design RELIC, an interactive system that enables users to investigate and verify semantic-level variations in multiple long-form responses. This allows users to recognize potentially inaccurate information in the generated text and make necessary corrections. From a user study with ten participants, we demonstrate that our approach helps users better verify the reliability of the generated text. We further summarize the design implications and lessons learned from this research for future studies of reliable human-LLM interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2311_16842
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle RELIC: Investigating Large Language Model Responses using Self-Consistency
Cheng, Furui
Zouhar, Vilém
Arora, Simran
Sachan, Mrinmaya
Strobelt, Hendrik
El-Assady, Mennatallah
Human-Computer Interaction
Computation and Language
Large Language Models (LLMs) are notorious for blending fact with fiction and generating non-factual content, known as hallucinations. To address this challenge, we propose an interactive system that helps users gain insight into the reliability of the generated text. Our approach is based on the idea that the self-consistency of multiple samples generated by the same LLM relates to its confidence in individual claims in the generated texts. Using this idea, we design RELIC, an interactive system that enables users to investigate and verify semantic-level variations in multiple long-form responses. This allows users to recognize potentially inaccurate information in the generated text and make necessary corrections. From a user study with ten participants, we demonstrate that our approach helps users better verify the reliability of the generated text. We further summarize the design implications and lessons learned from this research for future studies of reliable human-LLM interactions.
title RELIC: Investigating Large Language Model Responses using Self-Consistency
topic Human-Computer Interaction
Computation and Language
url https://arxiv.org/abs/2311.16842