Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Smit, Andries, Duckworth, Paul, Grinsztajn, Nathan, Barrett, Thomas D., Pretorius, Arnu
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917725942579200
author Smit, Andries
Duckworth, Paul
Grinsztajn, Nathan
Barrett, Thomas D.
Pretorius, Arnu
author_facet Smit, Andries
Duckworth, Paul
Grinsztajn, Nathan
Barrett, Thomas D.
Pretorius, Arnu
contents Recent advancements in large language models (LLMs) underscore their potential for responding to inquiries in various domains. However, ensuring that generative agents provide accurate and reliable answers remains an ongoing challenge. In this context, multi-agent debate (MAD) has emerged as a promising strategy for enhancing the truthfulness of LLMs. We benchmark a range of debating and prompting strategies to explore the trade-offs between cost, time, and accuracy. Importantly, we find that multi-agent debating systems, in their current form, do not reliably outperform other proposed prompting strategies, such as self-consistency and ensembling using multiple reasoning paths. However, when performing hyperparameter tuning, several MAD systems, such as Multi-Persona, perform better. This suggests that MAD protocols might not be inherently worse than other approaches, but that they are more sensitive to different hyperparameter settings and difficult to optimize. We build on these results to offer insights into improving debating strategies, such as adjusting agent agreement levels, which can significantly enhance performance and even surpass all other non-debate protocols we evaluated. We provide an open-source repository to the community with several state-of-the-art protocols together with evaluation scripts to benchmark across popular research datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2311_17371
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs
Smit, Andries
Duckworth, Paul
Grinsztajn, Nathan
Barrett, Thomas D.
Pretorius, Arnu
Computation and Language
Artificial Intelligence
Recent advancements in large language models (LLMs) underscore their potential for responding to inquiries in various domains. However, ensuring that generative agents provide accurate and reliable answers remains an ongoing challenge. In this context, multi-agent debate (MAD) has emerged as a promising strategy for enhancing the truthfulness of LLMs. We benchmark a range of debating and prompting strategies to explore the trade-offs between cost, time, and accuracy. Importantly, we find that multi-agent debating systems, in their current form, do not reliably outperform other proposed prompting strategies, such as self-consistency and ensembling using multiple reasoning paths. However, when performing hyperparameter tuning, several MAD systems, such as Multi-Persona, perform better. This suggests that MAD protocols might not be inherently worse than other approaches, but that they are more sensitive to different hyperparameter settings and difficult to optimize. We build on these results to offer insights into improving debating strategies, such as adjusting agent agreement levels, which can significantly enhance performance and even surpass all other non-debate protocols we evaluated. We provide an open-source repository to the community with several state-of-the-art protocols together with evaluation scripts to benchmark across popular research datasets.
title Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2311.17371