Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Subramaniam, Vighnesh, Du, Yilun, Tenenbaum, Joshua B., Torralba, Antonio, Li, Shuang, Mordatch, Igor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917943793680384
author Subramaniam, Vighnesh
Du, Yilun
Tenenbaum, Joshua B.
Torralba, Antonio
Li, Shuang
Mordatch, Igor
author_facet Subramaniam, Vighnesh
Du, Yilun
Tenenbaum, Joshua B.
Torralba, Antonio
Li, Shuang
Mordatch, Igor
contents Large language models (LLMs) have achieved remarkable performance in recent years but are fundamentally limited by the underlying training data. To improve models beyond the training data, recent works have explored how LLMs can be used to generate synthetic data for autonomous self-improvement. However, successive steps of self-improvement can reach a point of diminishing returns. In this work, we propose a complementary approach towards self-improvement where finetuning is applied to a multiagent society of language models. A group of language models, all starting from the same base model, are independently specialized by updating each one using data generated through multiagent interactions among the models. By training each model on independent sets of data, we illustrate how this approach enables specialization across models and diversification over the set of models. As a result, our overall system is able to preserve diverse reasoning chains and autonomously improve over many more rounds of fine-tuning than single-agent self-improvement methods. We quantitatively illustrate the efficacy of the approach across a wide suite of reasoning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2501_05707
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains
Subramaniam, Vighnesh
Du, Yilun
Tenenbaum, Joshua B.
Torralba, Antonio
Li, Shuang
Mordatch, Igor
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) have achieved remarkable performance in recent years but are fundamentally limited by the underlying training data. To improve models beyond the training data, recent works have explored how LLMs can be used to generate synthetic data for autonomous self-improvement. However, successive steps of self-improvement can reach a point of diminishing returns. In this work, we propose a complementary approach towards self-improvement where finetuning is applied to a multiagent society of language models. A group of language models, all starting from the same base model, are independently specialized by updating each one using data generated through multiagent interactions among the models. By training each model on independent sets of data, we illustrate how this approach enables specialization across models and diversification over the set of models. As a result, our overall system is able to preserve diverse reasoning chains and autonomously improve over many more rounds of fine-tuning than single-agent self-improvement methods. We quantitatively illustrate the efficacy of the approach across a wide suite of reasoning tasks.
title Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2501.05707