MALLM: Multi-Agent Large Language Models Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Becker, Jonas, Kaesberg, Lars Benedikt, Bauer, Niklas, Wahle, Jan Philip, Ruas, Terry, Gipp, Bela
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918248670298112
author Becker, Jonas
Kaesberg, Lars Benedikt
Bauer, Niklas
Wahle, Jan Philip
Ruas, Terry
Gipp, Bela
author_facet Becker, Jonas
Kaesberg, Lars Benedikt
Bauer, Niklas
Wahle, Jan Philip
Ruas, Terry
Gipp, Bela
contents Multi-agent debate (MAD) has demonstrated the ability to augment collective intelligence by scaling test-time compute and leveraging expertise. Current frameworks for multi-agent debate are often designed towards tool use, lack integrated evaluation, or provide limited configurability of agent personas, response generators, discussion paradigms, and decision protocols. We introduce MALLM (Multi-Agent Large Language Models), an open-source framework that enables systematic analysis of MAD components. MALLM offers more than 144 unique configurations of MAD, including (1) agent personas (e.g., Expert, Personality), (2) response generators (e.g., Critical, Reasoning), (3) discussion paradigms (e.g., Memory, Relay), and (4) decision protocols (e.g., Voting, Consensus). MALLM uses simple configuration files to define a debate. Furthermore, MALLM can load any textual Hugging Face dataset (e.g., MMLU-Pro, WinoGrande) and provides an evaluation pipeline for easy comparison of MAD configurations. MALLM enables researchers to systematically configure, run, and evaluate debates for their problems, facilitating the understanding of the components and their interplay.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11656
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MALLM: Multi-Agent Large Language Models Framework
Becker, Jonas
Kaesberg, Lars Benedikt
Bauer, Niklas
Wahle, Jan Philip
Ruas, Terry
Gipp, Bela
Multiagent Systems
Artificial Intelligence
Computation and Language
A.1; I.2.7
Multi-agent debate (MAD) has demonstrated the ability to augment collective intelligence by scaling test-time compute and leveraging expertise. Current frameworks for multi-agent debate are often designed towards tool use, lack integrated evaluation, or provide limited configurability of agent personas, response generators, discussion paradigms, and decision protocols. We introduce MALLM (Multi-Agent Large Language Models), an open-source framework that enables systematic analysis of MAD components. MALLM offers more than 144 unique configurations of MAD, including (1) agent personas (e.g., Expert, Personality), (2) response generators (e.g., Critical, Reasoning), (3) discussion paradigms (e.g., Memory, Relay), and (4) decision protocols (e.g., Voting, Consensus). MALLM uses simple configuration files to define a debate. Furthermore, MALLM can load any textual Hugging Face dataset (e.g., MMLU-Pro, WinoGrande) and provides an evaluation pipeline for easy comparison of MAD configurations. MALLM enables researchers to systematically configure, run, and evaluate debates for their problems, facilitating the understanding of the components and their interplay.
title MALLM: Multi-Agent Large Language Models Framework
topic Multiagent Systems
Artificial Intelligence
Computation and Language
A.1; I.2.7
url https://arxiv.org/abs/2509.11656