When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rykov, Elisei, Petrushina, Kseniia, Savkin, Maksim, Olisov, Valerii, Vazhentsev, Artem, Titova, Kseniia, Panchenko, Alexander, Konovalov, Vasily, Belikova, Julia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914077654122496
author Rykov, Elisei
Petrushina, Kseniia
Savkin, Maksim
Olisov, Valerii
Vazhentsev, Artem
Titova, Kseniia
Panchenko, Alexander
Konovalov, Vasily
Belikova, Julia
author_facet Rykov, Elisei
Petrushina, Kseniia
Savkin, Maksim
Olisov, Valerii
Vazhentsev, Artem
Titova, Kseniia
Panchenko, Alexander
Konovalov, Vasily
Belikova, Julia
contents Hallucination detection remains a fundamental challenge for the safe and reliable deployment of large language models (LLMs), especially in applications requiring factual accuracy. Existing hallucination benchmarks often operate at the sequence level and are limited to English, lacking the fine-grained, multilingual supervision needed for a comprehensive evaluation. In this work, we introduce PsiloQA, a large-scale, multilingual dataset annotated with span-level hallucinations across 14 languages. PsiloQA is constructed through an automated three-stage pipeline: generating question-answer pairs from Wikipedia using GPT-4o, eliciting potentially hallucinated answers from diverse LLMs in a no-context setting, and automatically annotating hallucinated spans using GPT-4o by comparing against golden answers and retrieved context. We evaluate a wide range of hallucination detection methods -- including uncertainty quantification, LLM-based tagging, and fine-tuned encoder models -- and show that encoder-based models achieve the strongest performance across languages. Furthermore, PsiloQA demonstrates effective cross-lingual generalization and supports robust knowledge transfer to other benchmarks, all while being significantly more cost-efficient than human-annotated datasets. Our dataset and results advance the development of scalable, fine-grained hallucination detection in multilingual settings.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04849
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
Rykov, Elisei
Petrushina, Kseniia
Savkin, Maksim
Olisov, Valerii
Vazhentsev, Artem
Titova, Kseniia
Panchenko, Alexander
Konovalov, Vasily
Belikova, Julia
Computation and Language
Hallucination detection remains a fundamental challenge for the safe and reliable deployment of large language models (LLMs), especially in applications requiring factual accuracy. Existing hallucination benchmarks often operate at the sequence level and are limited to English, lacking the fine-grained, multilingual supervision needed for a comprehensive evaluation. In this work, we introduce PsiloQA, a large-scale, multilingual dataset annotated with span-level hallucinations across 14 languages. PsiloQA is constructed through an automated three-stage pipeline: generating question-answer pairs from Wikipedia using GPT-4o, eliciting potentially hallucinated answers from diverse LLMs in a no-context setting, and automatically annotating hallucinated spans using GPT-4o by comparing against golden answers and retrieved context. We evaluate a wide range of hallucination detection methods -- including uncertainty quantification, LLM-based tagging, and fine-tuned encoder models -- and show that encoder-based models achieve the strongest performance across languages. Furthermore, PsiloQA demonstrates effective cross-lingual generalization and supports robust knowledge transfer to other benchmarks, all while being significantly more cost-efficient than human-annotated datasets. Our dataset and results advance the development of scalable, fine-grained hallucination detection in multilingual settings.
title When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
topic Computation and Language
url https://arxiv.org/abs/2510.04849