Liars' Bench: Evaluating Lie Detectors for Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kretschmar, Kieron, Laurito, Walter, Maiya, Sharan, Marks, Samuel
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908754511921152
author Kretschmar, Kieron
Laurito, Walter
Maiya, Sharan
Marks, Samuel
author_facet Kretschmar, Kieron
Laurito, Walter
Maiya, Sharan
Marks, Samuel
contents Prior work has introduced techniques for detecting when large language models (LLMs) lie, that is, generate statements they believe are false. However, these techniques are typically validated in narrow settings that do not capture the diverse lies LLMs can generate. We introduce LIARS' BENCH, a testbed consisting of 72,863 examples of lies and honest responses generated by four open-weight models across seven datasets. Our settings capture qualitatively different types of lies and vary along two dimensions: the model's reason for lying and the object of belief targeted by the lie. Evaluating three black- and white-box lie detection techniques on LIARS' BENCH, we find that existing techniques systematically fail to identify certain types of lies, especially in settings where it's not possible to determine whether the model lied from the transcript alone. Overall, LIARS' BENCH reveals limitations in prior techniques and provides a practical testbed for guiding progress in lie detection.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16035
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Liars' Bench: Evaluating Lie Detectors for Language Models
Kretschmar, Kieron
Laurito, Walter
Maiya, Sharan
Marks, Samuel
Computation and Language
Artificial Intelligence
Prior work has introduced techniques for detecting when large language models (LLMs) lie, that is, generate statements they believe are false. However, these techniques are typically validated in narrow settings that do not capture the diverse lies LLMs can generate. We introduce LIARS' BENCH, a testbed consisting of 72,863 examples of lies and honest responses generated by four open-weight models across seven datasets. Our settings capture qualitatively different types of lies and vary along two dimensions: the model's reason for lying and the object of belief targeted by the lie. Evaluating three black- and white-box lie detection techniques on LIARS' BENCH, we find that existing techniques systematically fail to identify certain types of lies, especially in settings where it's not possible to determine whether the model lied from the transcript alone. Overall, LIARS' BENCH reveals limitations in prior techniques and provides a practical testbed for guiding progress in lie detection.
title Liars' Bench: Evaluating Lie Detectors for Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.16035