Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jhaveri, Ayush Rajesh, GX-Chen, Anthony, Sucholutsky, Ilia, Choi, Eunsol
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913000784396288
author Jhaveri, Ayush Rajesh
GX-Chen, Anthony
Sucholutsky, Ilia
Choi, Eunsol
author_facet Jhaveri, Ayush Rajesh
GX-Chen, Anthony
Sucholutsky, Ilia
Choi, Eunsol
contents Confirmation bias, the tendency to seek evidence that supports rather than challenges one's belief, hinders one's reasoning ability. We examine whether large language models (LLMs) exhibit confirmation bias by adapting the rule-discovery study from human psychology: given a sequence of three numbers (a "triple"), an agent engages in an interactive feedback loop where it (1) proposes a new triple, (2) receives feedback on whether it satisfies the hidden rule, and (3) guesses the rule. Across eleven LLMs of multiple families and scales, we find that LLMs exhibit confirmation bias, often proposing triples to confirm their hypothesis rather than trying to falsify it. This leads to slower and less frequent discovery of the hidden rule. We further explore intervention strategies (e.g., encouraging the agent to consider counter examples) developed for humans. We find prompting LLMs with such instruction consistently decreases confirmation bias in LLMs, improving rule discovery rates from 42% to 56% on average. Lastly, we mitigate confirmation bias by distilling intervention-induced behavior into LLMs, showing promising generalization to a new task, the Blicket test. Our work shows that confirmation bias is a limitation of LLMs in hypothesis exploration, and that it can be mitigated via injecting interventions designed for humans.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02485
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
Jhaveri, Ayush Rajesh
GX-Chen, Anthony
Sucholutsky, Ilia
Choi, Eunsol
Computation and Language
Machine Learning
Confirmation bias, the tendency to seek evidence that supports rather than challenges one's belief, hinders one's reasoning ability. We examine whether large language models (LLMs) exhibit confirmation bias by adapting the rule-discovery study from human psychology: given a sequence of three numbers (a "triple"), an agent engages in an interactive feedback loop where it (1) proposes a new triple, (2) receives feedback on whether it satisfies the hidden rule, and (3) guesses the rule. Across eleven LLMs of multiple families and scales, we find that LLMs exhibit confirmation bias, often proposing triples to confirm their hypothesis rather than trying to falsify it. This leads to slower and less frequent discovery of the hidden rule. We further explore intervention strategies (e.g., encouraging the agent to consider counter examples) developed for humans. We find prompting LLMs with such instruction consistently decreases confirmation bias in LLMs, improving rule discovery rates from 42% to 56% on average. Lastly, we mitigate confirmation bias by distilling intervention-induced behavior into LLMs, showing promising generalization to a new task, the Blicket test. Our work shows that confirmation bias is a limitation of LLMs in hypothesis exploration, and that it can be mitigated via injecting interventions designed for humans.
title Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2604.02485