Can LLMs Detect Intrinsic Hallucinations in Paraphrasing and Machine Translation?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gogoulou, Evangelia, Zahra, Shorouq, Guillou, Liane, Dürlich, Luise, Nivre, Joakim
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908343097884672
author Gogoulou, Evangelia
Zahra, Shorouq
Guillou, Liane
Dürlich, Luise
Nivre, Joakim
author_facet Gogoulou, Evangelia
Zahra, Shorouq
Guillou, Liane
Dürlich, Luise
Nivre, Joakim
contents A frequently observed problem with LLMs is their tendency to generate output that is nonsensical, illogical, or factually incorrect, often referred to broadly as hallucination. Building on the recently proposed HalluciGen task for hallucination detection and generation, we evaluate a suite of open-access LLMs on their ability to detect intrinsic hallucinations in two conditional generation tasks: translation and paraphrasing. We study how model performance varies across tasks and language and we investigate the impact of model size, instruction tuning, and prompt choice. We find that performance varies across models but is consistent across prompts. Finally, we find that NLI models perform comparably well, suggesting that LLM-based detectors are not the only viable option for this specific task.
format Preprint
id arxiv_https___arxiv_org_abs_2504_20699
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can LLMs Detect Intrinsic Hallucinations in Paraphrasing and Machine Translation?
Gogoulou, Evangelia
Zahra, Shorouq
Guillou, Liane
Dürlich, Luise
Nivre, Joakim
Computation and Language
Artificial Intelligence
A frequently observed problem with LLMs is their tendency to generate output that is nonsensical, illogical, or factually incorrect, often referred to broadly as hallucination. Building on the recently proposed HalluciGen task for hallucination detection and generation, we evaluate a suite of open-access LLMs on their ability to detect intrinsic hallucinations in two conditional generation tasks: translation and paraphrasing. We study how model performance varies across tasks and language and we investigate the impact of model size, instruction tuning, and prompt choice. We find that performance varies across models but is consistent across prompts. Finally, we find that NLI models perform comparably well, suggesting that LLM-based detectors are not the only viable option for this specific task.
title Can LLMs Detect Intrinsic Hallucinations in Paraphrasing and Machine Translation?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.20699