Large Language Models Are Still Misled by Simple Bias Ensembles

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Zhouhao, Kan, Zhiyuan, Ding, Xiao, Du, Li, Cai, Bibo, Zhao, Yang, Qin, Bing, Liu, Ting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908976778575872
author Sun, Zhouhao
Kan, Zhiyuan
Ding, Xiao
Du, Li
Cai, Bibo
Zhao, Yang
Qin, Bing
Liu, Ting
author_facet Sun, Zhouhao
Kan, Zhiyuan
Ding, Xiao
Du, Li
Cai, Bibo
Zhao, Yang
Qin, Bing
Liu, Ting
contents With the evolution of large language models (LLMs), their robustness against individual simple biases has been enhanced. However, we observe that the ensemble of multiple simple biases still exerts a significant adverse impact on LLMs. Given that real-world data samples are typically confounded by a wide range of biases, LLMs tend to exhibit unstable performance when deployed in high-stakes real-world scenarios such as clinical diagnosis and legal document analysis. However, previous benchmarks are constrained to datasets where each sample is manually injected with only one type of bias. To bridge this gap, we propose a multi-bias benchmark where each sample contains multiple types of biases. Experimental results reveal that existing LLMs and debiasing methods perform poorly on this benchmark, highlighting the challenge of eliminating such compounded biases.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16522
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Models Are Still Misled by Simple Bias Ensembles
Sun, Zhouhao
Kan, Zhiyuan
Ding, Xiao
Du, Li
Cai, Bibo
Zhao, Yang
Qin, Bing
Liu, Ting
Computation and Language
Artificial Intelligence
With the evolution of large language models (LLMs), their robustness against individual simple biases has been enhanced. However, we observe that the ensemble of multiple simple biases still exerts a significant adverse impact on LLMs. Given that real-world data samples are typically confounded by a wide range of biases, LLMs tend to exhibit unstable performance when deployed in high-stakes real-world scenarios such as clinical diagnosis and legal document analysis. However, previous benchmarks are constrained to datasets where each sample is manually injected with only one type of bias. To bridge this gap, we propose a multi-bias benchmark where each sample contains multiple types of biases. Experimental results reveal that existing LLMs and debiasing methods perform poorly on this benchmark, highlighting the challenge of eliminating such compounded biases.
title Large Language Models Are Still Misled by Simple Bias Ensembles
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.16522