Auditing Multi-Agent LLM Reasoning Trees Outperforms Majority Vote and LLM-as-Judge

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Wei, Li, Shixuan, Ping, Heng, Zhang, Peiyu, Bogdan, Paul, Thomason, Jesse
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917263191310336
author Yang, Wei
Li, Shixuan
Ping, Heng
Zhang, Peiyu
Bogdan, Paul
Thomason, Jesse
author_facet Yang, Wei
Li, Shixuan
Ping, Heng
Zhang, Peiyu
Bogdan, Paul
Thomason, Jesse
contents Multi-agent systems (MAS) can substantially extend the reasoning capacity of large language models (LLMs), yet most frameworks still aggregate agent outputs with majority voting. This heuristic discards the evidential structure of reasoning traces and is brittle under the confabulation consensus, where agents share correlated biases and converge on the same incorrect rationale. We introduce AgentAuditor, which replaces voting with a path search over a Reasoning Tree that explicitly represents agreements and divergences among agent traces. AgentAuditor resolves conflicts by comparing reasoning branches at critical divergence points, turning global adjudication into efficient, localized verification. We further propose Anti-Consensus Preference Optimization (ACPO), which trains the adjudicator on majority-failure cases and rewards evidence-based minority selections over popular errors. AgentAuditor is agnostic to MAS setting, and we find across 5 popular settings that it yields up to 5% absolute accuracy improvement over a majority vote, and up to 3% over using LLM-as-Judge.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09341
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Auditing Multi-Agent LLM Reasoning Trees Outperforms Majority Vote and LLM-as-Judge
Yang, Wei
Li, Shixuan
Ping, Heng
Zhang, Peiyu
Bogdan, Paul
Thomason, Jesse
Artificial Intelligence
Multi-agent systems (MAS) can substantially extend the reasoning capacity of large language models (LLMs), yet most frameworks still aggregate agent outputs with majority voting. This heuristic discards the evidential structure of reasoning traces and is brittle under the confabulation consensus, where agents share correlated biases and converge on the same incorrect rationale. We introduce AgentAuditor, which replaces voting with a path search over a Reasoning Tree that explicitly represents agreements and divergences among agent traces. AgentAuditor resolves conflicts by comparing reasoning branches at critical divergence points, turning global adjudication into efficient, localized verification. We further propose Anti-Consensus Preference Optimization (ACPO), which trains the adjudicator on majority-failure cases and rewards evidence-based minority selections over popular errors. AgentAuditor is agnostic to MAS setting, and we find across 5 popular settings that it yields up to 5% absolute accuracy improvement over a majority vote, and up to 3% over using LLM-as-Judge.
title Auditing Multi-Agent LLM Reasoning Trees Outperforms Majority Vote and LLM-as-Judge
topic Artificial Intelligence
url https://arxiv.org/abs/2602.09341