Making Bias Non-Predictive: Training Robust LLM Reasoning via Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Qian, Zhao, Xuandong, Zhang, Zirui, Lou, Zhanzhi, Chen, Nuo, Song, Dawn, He, Bingsheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913006509621248
author Wang, Qian
Zhao, Xuandong
Zhang, Zirui
Lou, Zhanzhi
Chen, Nuo
Song, Dawn
He, Bingsheng
author_facet Wang, Qian
Zhao, Xuandong
Zhang, Zirui
Lou, Zhanzhi
Chen, Nuo
Song, Dawn
He, Bingsheng
contents Large language models (LLMs) increasingly serve as reasoners and automated evaluators, yet they remain susceptible to cognitive biases -- often altering their reasoning when faced with spurious prompt-level cues such as consensus claims or authority appeals.} Existing mitigations via prompting or supervised fine-tuning fail to generalize, as they modify surface behavior without changing the optimization objective that makes bias cues attractive. We propose \textbf{Epistemic Independence Training (EIT)}, a reinforcement learning framework grounded in a key principle: to learn independence, bias cues must be made non-predictive of reward. EIT operationalizes this through a balanced conflict strategy where bias signals are equally likely to support correct and incorrect answers, combined with a reward design that penalizes bias-following without rewarding bias agreement. Experiments on Qwen3-4B demonstrate that EIT improves both accuracy and robustness under adversarial biases, while preserving performance when bias aligns with truth. Notably, models trained only on bandwagon bias generalize to unseen bias types such as authority and distraction, indicating that EIT induces transferable epistemic independence rather than bias-specific heuristics. \revised{EIT further generalizes across benchmarks (MedQA, HellaSwag), model families (Llama-3.2-3B), and scales (Qwen3-8B), and outperforms distribution-shift methods (GroupDRO, IRM) without requiring environment labels.} Code and data are available at https://anonymous.4open.science/r/bias-mitigation-with-rl-BC47
format Preprint
id arxiv_https___arxiv_org_abs_2602_01528
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Making Bias Non-Predictive: Training Robust LLM Reasoning via Reinforcement Learning
Wang, Qian
Zhao, Xuandong
Zhang, Zirui
Lou, Zhanzhi
Chen, Nuo
Song, Dawn
He, Bingsheng
Computers and Society
Machine Learning
Large language models (LLMs) increasingly serve as reasoners and automated evaluators, yet they remain susceptible to cognitive biases -- often altering their reasoning when faced with spurious prompt-level cues such as consensus claims or authority appeals.} Existing mitigations via prompting or supervised fine-tuning fail to generalize, as they modify surface behavior without changing the optimization objective that makes bias cues attractive. We propose \textbf{Epistemic Independence Training (EIT)}, a reinforcement learning framework grounded in a key principle: to learn independence, bias cues must be made non-predictive of reward. EIT operationalizes this through a balanced conflict strategy where bias signals are equally likely to support correct and incorrect answers, combined with a reward design that penalizes bias-following without rewarding bias agreement. Experiments on Qwen3-4B demonstrate that EIT improves both accuracy and robustness under adversarial biases, while preserving performance when bias aligns with truth. Notably, models trained only on bandwagon bias generalize to unseen bias types such as authority and distraction, indicating that EIT induces transferable epistemic independence rather than bias-specific heuristics. \revised{EIT further generalizes across benchmarks (MedQA, HellaSwag), model families (Llama-3.2-3B), and scales (Qwen3-8B), and outperforms distribution-shift methods (GroupDRO, IRM) without requiring environment labels.} Code and data are available at https://anonymous.4open.science/r/bias-mitigation-with-rl-BC47
title Making Bias Non-Predictive: Training Robust LLM Reasoning via Reinforcement Learning
topic Computers and Society
Machine Learning
url https://arxiv.org/abs/2602.01528