MAMM-Refine: A Recipe for Improving Faithfulness in Generation with Multi-Agent Collaboration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wan, David, Chen, Justin Chih-Yao, Stengel-Eskin, Elias, Bansal, Mohit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910884110008320
author Wan, David
Chen, Justin Chih-Yao
Stengel-Eskin, Elias
Bansal, Mohit
author_facet Wan, David
Chen, Justin Chih-Yao
Stengel-Eskin, Elias
Bansal, Mohit
contents Multi-agent collaboration among models has shown promise in reasoning tasks but is underexplored in long-form generation tasks like summarization and question-answering. We extend multi-agent multi-model reasoning to generation, specifically to improving faithfulness through refinement, i.e., revising model-generated outputs to remove factual inconsistencies. We investigate how iterative collaboration among multiple instances and types of large language models (LLMs) enhances subtasks in the refinement process, such as error detection, critiquing unfaithful sentences, and making corrections based on critiques. We design intrinsic evaluations for each subtask, with our findings indicating that both multi-agent (multiple instances) and multi-model (diverse LLM types) approaches benefit error detection and critiquing. Additionally, reframing critiquing and refinement as reranking rather than generation tasks improves multi-agent performance. We consolidate these insights into a final "recipe" called Multi-Agent Multi-Model Refinement (MAMM-Refine), where multi-agent and multi-model collaboration significantly boosts performance on three summarization datasets as well as on long-form question answering, demonstrating the effectiveness and generalizability of our recipe.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15272
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MAMM-Refine: A Recipe for Improving Faithfulness in Generation with Multi-Agent Collaboration
Wan, David
Chen, Justin Chih-Yao
Stengel-Eskin, Elias
Bansal, Mohit
Computation and Language
Artificial Intelligence
Multi-agent collaboration among models has shown promise in reasoning tasks but is underexplored in long-form generation tasks like summarization and question-answering. We extend multi-agent multi-model reasoning to generation, specifically to improving faithfulness through refinement, i.e., revising model-generated outputs to remove factual inconsistencies. We investigate how iterative collaboration among multiple instances and types of large language models (LLMs) enhances subtasks in the refinement process, such as error detection, critiquing unfaithful sentences, and making corrections based on critiques. We design intrinsic evaluations for each subtask, with our findings indicating that both multi-agent (multiple instances) and multi-model (diverse LLM types) approaches benefit error detection and critiquing. Additionally, reframing critiquing and refinement as reranking rather than generation tasks improves multi-agent performance. We consolidate these insights into a final "recipe" called Multi-Agent Multi-Model Refinement (MAMM-Refine), where multi-agent and multi-model collaboration significantly boosts performance on three summarization datasets as well as on long-form question answering, demonstrating the effectiveness and generalizability of our recipe.
title MAMM-Refine: A Recipe for Improving Faithfulness in Generation with Multi-Agent Collaboration
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.15272