Do LLM hallucination detectors suffer from low-resource effect?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Datta, Debtanu, Chilukuri, Mohan Kishore, Kumar, Yash, Ghosh, Saptarshi, Zafar, Muhammad Bilal
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909999011201024
author Datta, Debtanu
Chilukuri, Mohan Kishore
Kumar, Yash
Ghosh, Saptarshi
Zafar, Muhammad Bilal
author_facet Datta, Debtanu
Chilukuri, Mohan Kishore
Kumar, Yash
Ghosh, Saptarshi
Zafar, Muhammad Bilal
contents LLMs, while outperforming humans in a wide range of tasks, can still fail in unanticipated ways. We focus on two pervasive failure modes: (i) hallucinations, where models produce incorrect information about the world, and (ii) the low-resource effect, where the models show impressive performance in high-resource languages like English but the performance degrades significantly in low-resource languages like Bengali. We study the intersection of these issues and ask: do hallucination detectors suffer from the low-resource effect? We conduct experiments on five tasks across three domains (factual recall, STEM, and Humanities). Experiments with four LLMs and three hallucination detectors reveal a curious finding: As expected, the task accuracies in low-resource languages experience large drops (compared to English). However, the drop in detectors' accuracy is often several times smaller than the drop in task accuracy. Our findings suggest that even in low-resource languages, the internal mechanisms of LLMs might encode signals about their uncertainty. Further, the detectors are robust within language (even for non-English) and in multilingual setups, but not in cross-lingual settings without in-language supervision.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16766
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Do LLM hallucination detectors suffer from low-resource effect?
Datta, Debtanu
Chilukuri, Mohan Kishore
Kumar, Yash
Ghosh, Saptarshi
Zafar, Muhammad Bilal
Computation and Language
Artificial Intelligence
LLMs, while outperforming humans in a wide range of tasks, can still fail in unanticipated ways. We focus on two pervasive failure modes: (i) hallucinations, where models produce incorrect information about the world, and (ii) the low-resource effect, where the models show impressive performance in high-resource languages like English but the performance degrades significantly in low-resource languages like Bengali. We study the intersection of these issues and ask: do hallucination detectors suffer from the low-resource effect? We conduct experiments on five tasks across three domains (factual recall, STEM, and Humanities). Experiments with four LLMs and three hallucination detectors reveal a curious finding: As expected, the task accuracies in low-resource languages experience large drops (compared to English). However, the drop in detectors' accuracy is often several times smaller than the drop in task accuracy. Our findings suggest that even in low-resource languages, the internal mechanisms of LLMs might encode signals about their uncertainty. Further, the detectors are robust within language (even for non-English) and in multilingual setups, but not in cross-lingual settings without in-language supervision.
title Do LLM hallucination detectors suffer from low-resource effect?
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.16766