Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wynn, Andrea, Satija, Harsh, Hadfield, Gillian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914090122739712
author Wynn, Andrea
Satija, Harsh
Hadfield, Gillian
author_facet Wynn, Andrea
Satija, Harsh
Hadfield, Gillian
contents While multi-agent debate has been proposed as a promising strategy for improving AI reasoning ability, we find that debate can sometimes be harmful rather than helpful. Prior work has primarily focused on debates within homogeneous groups of agents, whereas we explore how diversity in model capabilities influences the dynamics and outcomes of multi-agent interactions. Through a series of experiments, we demonstrate that debate can lead to a decrease in accuracy over time - even in settings where stronger (i.e., more capable) models outnumber their weaker counterparts. Our analysis reveals that models frequently shift from correct to incorrect answers in response to peer reasoning, favoring agreement over challenging flawed reasoning. We perform additional experiments investigating various potential contributing factors to these harmful shifts - including sycophancy, social conformity, and model and task type. These results highlight important failure modes in the exchange of reasons during multi-agent debate, suggesting that naive applications of debate may cause performance degradation when agents are neither incentivised nor adequately equipped to resist persuasive but incorrect reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05396
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate
Wynn, Andrea
Satija, Harsh
Hadfield, Gillian
Computation and Language
Artificial Intelligence
Multiagent Systems
While multi-agent debate has been proposed as a promising strategy for improving AI reasoning ability, we find that debate can sometimes be harmful rather than helpful. Prior work has primarily focused on debates within homogeneous groups of agents, whereas we explore how diversity in model capabilities influences the dynamics and outcomes of multi-agent interactions. Through a series of experiments, we demonstrate that debate can lead to a decrease in accuracy over time - even in settings where stronger (i.e., more capable) models outnumber their weaker counterparts. Our analysis reveals that models frequently shift from correct to incorrect answers in response to peer reasoning, favoring agreement over challenging flawed reasoning. We perform additional experiments investigating various potential contributing factors to these harmful shifts - including sycophancy, social conformity, and model and task type. These results highlight important failure modes in the exchange of reasons during multi-agent debate, suggesting that naive applications of debate may cause performance degradation when agents are neither incentivised nor adequately equipped to resist persuasive but incorrect reasoning.
title Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate
topic Computation and Language
Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2509.05396