Strong hallucinations from negation and how to fix them

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Asher, Nicholas, Bhar, Swarnadeep
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929464748802048
author Asher, Nicholas
Bhar, Swarnadeep
author_facet Asher, Nicholas
Bhar, Swarnadeep
contents Despite great performance on many tasks, language models (LMs) still struggle with reasoning, sometimes providing responses that cannot possibly be true because they stem from logical incoherence. We call such responses \textit{strong hallucinations} and prove that they follow from an LM's computation of its internal representations for logical operators and outputs from those representations. Focusing on negation, we provide a novel solution in which negation is treated not as another element of a latent representation, but as \textit{an operation over an LM's latent representations that constrains how they may evolve}. We show that our approach improves model performance in cloze prompting and natural language inference tasks with negation without requiring training on sparse negative data.
format Preprint
id arxiv_https___arxiv_org_abs_2402_10543
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Strong hallucinations from negation and how to fix them
Asher, Nicholas
Bhar, Swarnadeep
Computation and Language
Artificial Intelligence
I.2.7
Despite great performance on many tasks, language models (LMs) still struggle with reasoning, sometimes providing responses that cannot possibly be true because they stem from logical incoherence. We call such responses \textit{strong hallucinations} and prove that they follow from an LM's computation of its internal representations for logical operators and outputs from those representations. Focusing on negation, we provide a novel solution in which negation is treated not as another element of a latent representation, but as \textit{an operation over an LM's latent representations that constrains how they may evolve}. We show that our approach improves model performance in cloze prompting and natural language inference tasks with negation without requiring training on sparse negative data.
title Strong hallucinations from negation and how to fix them
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2402.10543