PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hosseini, Mohammad, Hosseini, Kimia, Bali, Shayan, Zanjani, Zahra, Momtazi, Saeedeh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916970137387008
author Hosseini, Mohammad
Hosseini, Kimia
Bali, Shayan
Zanjani, Zahra
Momtazi, Saeedeh
author_facet Hosseini, Mohammad
Hosseini, Kimia
Bali, Shayan
Zanjani, Zahra
Momtazi, Saeedeh
contents Hallucination is a persistent issue affecting all large language Models (LLMs), particularly within low-resource languages such as Persian. PerHalluEval (Persian Hallucination Evaluation) is the first dynamic hallucination evaluation benchmark tailored for the Persian language. Our benchmark leverages a three-stage LLM-driven pipeline, augmented with human validation, to generate plausible answers and summaries regarding QA and summarization tasks, focusing on detecting extrinsic and intrinsic hallucinations. Moreover, we used the log probabilities of generated tokens to select the most believable hallucinated instances. In addition, we engaged human annotators to highlight Persian-specific contexts in the QA dataset in order to evaluate LLMs' performance on content specifically related to Persian culture. Our evaluation of 12 LLMs, including open- and closed-source models using PerHalluEval, revealed that the models generally struggle in detecting hallucinated Persian text. We showed that providing external knowledge, i.e., the original document for the summarization task, could mitigate hallucination partially. Furthermore, there was no significant difference in terms of hallucination when comparing LLMs specifically trained for Persian with others.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21104
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
Hosseini, Mohammad
Hosseini, Kimia
Bali, Shayan
Zanjani, Zahra
Momtazi, Saeedeh
Computation and Language
Hallucination is a persistent issue affecting all large language Models (LLMs), particularly within low-resource languages such as Persian. PerHalluEval (Persian Hallucination Evaluation) is the first dynamic hallucination evaluation benchmark tailored for the Persian language. Our benchmark leverages a three-stage LLM-driven pipeline, augmented with human validation, to generate plausible answers and summaries regarding QA and summarization tasks, focusing on detecting extrinsic and intrinsic hallucinations. Moreover, we used the log probabilities of generated tokens to select the most believable hallucinated instances. In addition, we engaged human annotators to highlight Persian-specific contexts in the QA dataset in order to evaluate LLMs' performance on content specifically related to Persian culture. Our evaluation of 12 LLMs, including open- and closed-source models using PerHalluEval, revealed that the models generally struggle in detecting hallucinated Persian text. We showed that providing external knowledge, i.e., the original document for the summarization task, could mitigate hallucination partially. Furthermore, there was no significant difference in terms of hallucination when comparing LLMs specifically trained for Persian with others.
title PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2509.21104