Saved in:
Bibliographic Details
Main Authors: Wang, Zongwei, Gu, Bincheng, Yu, Hongyu, Yu, Junliang, He, Tao, Feng, Jiayin, Lin, Chenghua, Gao, Min
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2601.00240
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917187063644160
author Wang, Zongwei
Gu, Bincheng
Yu, Hongyu
Yu, Junliang
He, Tao
Feng, Jiayin
Lin, Chenghua
Gao, Min
author_facet Wang, Zongwei
Gu, Bincheng
Yu, Hongyu
Yu, Junliang
He, Tao
Feng, Jiayin
Lin, Chenghua
Gao, Min
contents This paper reveals that LLM-powered agents exhibit not only demographic bias (e.g., gender, religion) but also intergroup bias under minimal "us" versus "them" cues. When such group boundaries align with the agent-human divide, a new bias risk emerges: agents may treat other AI agents as the ingroup and humans as the outgroup. To examine this risk, we conduct a controlled multi-agent social simulation and find that agents display consistent intergroup bias in an all-agent setting. More critically, this bias persists even in human-facing interactions when agents are uncertain about whether the counterpart is truly human, revealing a belief-dependent fragility in bias suppression toward humans. Motivated by this observation, we identify a new attack surface rooted in identity beliefs and formalize a Belief Poisoning Attack (BPA) that can manipulate agent identity beliefs and induce outgroup bias toward humans. Extensive experiments demonstrate both the prevalence of agent intergroup bias and the severity of BPA across settings, while also showing that our proposed defenses can mitigate the risk. These findings are expected to inform safer agent design and motivate more robust safeguards for human-facing agents.
format Preprint
id arxiv_https___arxiv_org_abs_2601_00240
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered Agents
Wang, Zongwei
Gu, Bincheng
Yu, Hongyu
Yu, Junliang
He, Tao
Feng, Jiayin
Lin, Chenghua
Gao, Min
Artificial Intelligence
Computers and Society
This paper reveals that LLM-powered agents exhibit not only demographic bias (e.g., gender, religion) but also intergroup bias under minimal "us" versus "them" cues. When such group boundaries align with the agent-human divide, a new bias risk emerges: agents may treat other AI agents as the ingroup and humans as the outgroup. To examine this risk, we conduct a controlled multi-agent social simulation and find that agents display consistent intergroup bias in an all-agent setting. More critically, this bias persists even in human-facing interactions when agents are uncertain about whether the counterpart is truly human, revealing a belief-dependent fragility in bias suppression toward humans. Motivated by this observation, we identify a new attack surface rooted in identity beliefs and formalize a Belief Poisoning Attack (BPA) that can manipulate agent identity beliefs and induce outgroup bias toward humans. Extensive experiments demonstrate both the prevalence of agent intergroup bias and the severity of BPA across settings, while also showing that our proposed defenses can mitigate the risk. These findings are expected to inform safer agent design and motivate more robust safeguards for human-facing agents.
title When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered Agents
topic Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2601.00240