C2PO: Diagnosing and Disentangling Bias Shortcuts in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Xuan, An, Bo, Gu, Tianlong, Chang, Liang, Hao, Fengrui, Yu, Peipeng, Zhao, Shuai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909977696796672
author Feng, Xuan
An, Bo
Gu, Tianlong
Chang, Liang
Hao, Fengrui
Yu, Peipeng
Zhao, Shuai
author_facet Feng, Xuan
An, Bo
Gu, Tianlong
Chang, Liang
Hao, Fengrui
Yu, Peipeng
Zhao, Shuai
contents Bias in Large Language Models (LLMs) poses significant risks to trustworthiness, manifesting primarily as stereotypical biases (e.g., gender or racial stereotypes) and structural biases (e.g., lexical overlap or position preferences). However, prior paradigms typically address these in isolation, often mitigating one at the expense of exacerbating the other. To address this, we conduct a systematic exploration of these reasoning failures and identify a primary inducement: the latent spurious feature correlations within the input that drive these erroneous reasoning shortcuts. Driven by these findings, we introduce Causal-Contrastive Preference Optimization (C2PO), a unified alignment framework designed to tackle these specific failures by simultaneously discovering and suppressing these correlations directly within the optimization process. Specifically, C2PO leverages causal counterfactual signals to isolate bias-inducing features from valid reasoning paths, and employs a fairness-sensitive preference update mechanism to dynamically evaluate logit-level contributions and suppress shortcut features. Extensive experiments across multiple benchmarks covering stereotypical bias (BBQ, Unqover), structural bias (MNLI, HANS, Chatbot, MT-Bench), out-of-domain fairness (StereoSet, WinoBias), and general utility (MMLU, GSM8K) demonstrate that C2PO effectively mitigates stereotypical and structural biases while preserving robust general reasoning capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2512_23430
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle C2PO: Diagnosing and Disentangling Bias Shortcuts in LLMs
Feng, Xuan
An, Bo
Gu, Tianlong
Chang, Liang
Hao, Fengrui
Yu, Peipeng
Zhao, Shuai
Computation and Language
Bias in Large Language Models (LLMs) poses significant risks to trustworthiness, manifesting primarily as stereotypical biases (e.g., gender or racial stereotypes) and structural biases (e.g., lexical overlap or position preferences). However, prior paradigms typically address these in isolation, often mitigating one at the expense of exacerbating the other. To address this, we conduct a systematic exploration of these reasoning failures and identify a primary inducement: the latent spurious feature correlations within the input that drive these erroneous reasoning shortcuts. Driven by these findings, we introduce Causal-Contrastive Preference Optimization (C2PO), a unified alignment framework designed to tackle these specific failures by simultaneously discovering and suppressing these correlations directly within the optimization process. Specifically, C2PO leverages causal counterfactual signals to isolate bias-inducing features from valid reasoning paths, and employs a fairness-sensitive preference update mechanism to dynamically evaluate logit-level contributions and suppress shortcut features. Extensive experiments across multiple benchmarks covering stereotypical bias (BBQ, Unqover), structural bias (MNLI, HANS, Chatbot, MT-Bench), out-of-domain fairness (StereoSet, WinoBias), and general utility (MMLU, GSM8K) demonstrate that C2PO effectively mitigates stereotypical and structural biases while preserving robust general reasoning capabilities.
title C2PO: Diagnosing and Disentangling Bias Shortcuts in LLMs
topic Computation and Language
url https://arxiv.org/abs/2512.23430