Challenging Multilingual LLMs: A New Taxonomy and Benchmark for Unraveling Hallucination in Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Xinwei, Liu, Heng, Zhou, Jiang, Zhao, Xiaohu, Xu, Linlong, Wang, Longyue, Luo, Weihua, Zhang, Kaifu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917047822188544
author Wu, Xinwei
Liu, Heng
Zhou, Jiang
Zhao, Xiaohu
Xu, Linlong
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
author_facet Wu, Xinwei
Liu, Heng
Zhou, Jiang
Zhao, Xiaohu
Xu, Linlong
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
contents Large Language Models (LLMs) have advanced machine translation but remain vulnerable to hallucinations. Unfortunately, existing MT benchmarks are not capable of exposing failures in multilingual LLMs. To disclose hallucination in multilingual LLMs, we introduce a diagnostic framework with a taxonomy that separates Instruction Detachment from Source Detachment. Guided by this taxonomy, we create HalloMTBench, a multilingual, human-verified benchmark across 11 English-to-X directions. We employed 4 frontier LLMs to generate candidates and scrutinize these candidates with an ensemble of LLM judges, and expert validation. In this way, we curate 5,435 high-quality instances. We have evaluated 17 LLMs on HalloMTBench. Results reveal distinct ``hallucination triggers'' -- unique failure patterns reflecting model scale, source length sensitivity, linguistic biases, and Reinforcement-Learning (RL) amplified language mixing. HalloMTBench offers a forward-looking testbed for diagnosing LLM translation failures. HalloMTBench is available in https://huggingface.co/collections/AIDC-AI/marco-mt.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24073
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Challenging Multilingual LLMs: A New Taxonomy and Benchmark for Unraveling Hallucination in Translation
Wu, Xinwei
Liu, Heng
Zhou, Jiang
Zhao, Xiaohu
Xu, Linlong
Wang, Longyue
Luo, Weihua
Zhang, Kaifu
Computation and Language
Large Language Models (LLMs) have advanced machine translation but remain vulnerable to hallucinations. Unfortunately, existing MT benchmarks are not capable of exposing failures in multilingual LLMs. To disclose hallucination in multilingual LLMs, we introduce a diagnostic framework with a taxonomy that separates Instruction Detachment from Source Detachment. Guided by this taxonomy, we create HalloMTBench, a multilingual, human-verified benchmark across 11 English-to-X directions. We employed 4 frontier LLMs to generate candidates and scrutinize these candidates with an ensemble of LLM judges, and expert validation. In this way, we curate 5,435 high-quality instances. We have evaluated 17 LLMs on HalloMTBench. Results reveal distinct ``hallucination triggers'' -- unique failure patterns reflecting model scale, source length sensitivity, linguistic biases, and Reinforcement-Learning (RL) amplified language mixing. HalloMTBench offers a forward-looking testbed for diagnosing LLM translation failures. HalloMTBench is available in https://huggingface.co/collections/AIDC-AI/marco-mt.
title Challenging Multilingual LLMs: A New Taxonomy and Benchmark for Unraveling Hallucination in Translation
topic Computation and Language
url https://arxiv.org/abs/2510.24073