When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Choi, Hyeong Kyu, Zhu, Xiaojin, Li, Sharon
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908950041985024
author Choi, Hyeong Kyu
Zhu, Xiaojin
Li, Sharon
author_facet Choi, Hyeong Kyu
Zhu, Xiaojin
Li, Sharon
contents Multi-agent debate (MAD) aims to improve large language model (LLM) reasoning by letting multiple agents exchange answers and then aggregate their opinions. Yet recent studies reveal that agents are not neutral: they are prone to identity-driven sycophancy and self-bias, uncritically adopting a peer's view or stubbornly adhering to their own prior output, undermining the reliability of debate. In this work, we present the first principled framework that joins sycophancy and self-bias to mitigate and quantify identity bias in MAD. First, we formalize the debate dynamics as an identity-weighted Bayesian update process. Second, we propose response anonymization: by removing identity markers from prompts, agents cannot distinguish "self" from "peer", which forces equal weights on agent identity, thereby reducing bias and improving trustworthiness. Third, we define the Identity Bias Coefficient (IBC), a principled bias metric that measures an agent's tendency to follow its peer versus itself. Empirical studies across multiple models and benchmarks confirm that identity bias is widespread, with sycophancy far more common than self-bias. Our findings highlight the need to ensure that MAD systems reason based on content rather than identity. Code is released in https://github.com/deeplearning-wisc/MAD-identity-bias.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07517
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning
Choi, Hyeong Kyu
Zhu, Xiaojin
Li, Sharon
Artificial Intelligence
Multiagent Systems
Multi-agent debate (MAD) aims to improve large language model (LLM) reasoning by letting multiple agents exchange answers and then aggregate their opinions. Yet recent studies reveal that agents are not neutral: they are prone to identity-driven sycophancy and self-bias, uncritically adopting a peer's view or stubbornly adhering to their own prior output, undermining the reliability of debate. In this work, we present the first principled framework that joins sycophancy and self-bias to mitigate and quantify identity bias in MAD. First, we formalize the debate dynamics as an identity-weighted Bayesian update process. Second, we propose response anonymization: by removing identity markers from prompts, agents cannot distinguish "self" from "peer", which forces equal weights on agent identity, thereby reducing bias and improving trustworthiness. Third, we define the Identity Bias Coefficient (IBC), a principled bias metric that measures an agent's tendency to follow its peer versus itself. Empirical studies across multiple models and benchmarks confirm that identity bias is widespread, with sycophancy far more common than self-bias. Our findings highlight the need to ensure that MAD systems reason based on content rather than identity. Code is released in https://github.com/deeplearning-wisc/MAD-identity-bias.
title When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning
topic Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2510.07517