Improving Model Alignment Through Collective Intelligence of Open-Source LLMS

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Junlin, Xie, Roy, Zhu, Shang, Wang, Jue, Athiwaratkun, Ben, Dhingra, Bhuwan, Song, Shuaiwen Leon, Zhang, Ce, Zou, James
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910929221844992
author Wang, Junlin
Xie, Roy
Zhu, Shang
Wang, Jue
Athiwaratkun, Ben
Dhingra, Bhuwan
Song, Shuaiwen Leon
Zhang, Ce
Zou, James
author_facet Wang, Junlin
Xie, Roy
Zhu, Shang
Wang, Jue
Athiwaratkun, Ben
Dhingra, Bhuwan
Song, Shuaiwen Leon
Zhang, Ce
Zou, James
contents Building helpful and harmless large language models (LLMs) requires effective model alignment approach based on human instructions and feedback, which necessitates high-quality human-labeled data. Constructing such datasets is often expensive and hard to scale, and may face potential limitations on diversity and generalization. To address these challenges, we introduce Mixture of Agents Alignment (MoAA), that leverages the collective strengths of various language models to provide high-quality data for model alignment. By employing MoAA, we enhance both supervised fine-tuning and preference optimization, leading to improved performance compared to using a single model alone to generate alignment data (e.g. using GPT-4o alone). Evaluation results show that our approach can improve win rate of LLaMA-3.1-8B-Instruct from 19.5 to 48.3 on Arena-Hard and from 22.33 to 57.23 on AlpacaEval2, highlighting a promising direction for model alignment through this new scalable and diverse synthetic data recipe. Furthermore, we demonstrate that MoAA enables a self-improvement pipeline, where models finetuned on MoA-generated data surpass their own initial capabilities, providing evidence that our approach can push the frontier of open-source LLMs without reliance on stronger external supervision. Data and code will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2505_03059
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Model Alignment Through Collective Intelligence of Open-Source LLMS
Wang, Junlin
Xie, Roy
Zhu, Shang
Wang, Jue
Athiwaratkun, Ben
Dhingra, Bhuwan
Song, Shuaiwen Leon
Zhang, Ce
Zou, James
Computation and Language
Building helpful and harmless large language models (LLMs) requires effective model alignment approach based on human instructions and feedback, which necessitates high-quality human-labeled data. Constructing such datasets is often expensive and hard to scale, and may face potential limitations on diversity and generalization. To address these challenges, we introduce Mixture of Agents Alignment (MoAA), that leverages the collective strengths of various language models to provide high-quality data for model alignment. By employing MoAA, we enhance both supervised fine-tuning and preference optimization, leading to improved performance compared to using a single model alone to generate alignment data (e.g. using GPT-4o alone). Evaluation results show that our approach can improve win rate of LLaMA-3.1-8B-Instruct from 19.5 to 48.3 on Arena-Hard and from 22.33 to 57.23 on AlpacaEval2, highlighting a promising direction for model alignment through this new scalable and diverse synthetic data recipe. Furthermore, we demonstrate that MoAA enables a self-improvement pipeline, where models finetuned on MoA-generated data surpass their own initial capabilities, providing evidence that our approach can push the frontier of open-source LLMs without reliance on stronger external supervision. Data and code will be released.
title Improving Model Alignment Through Collective Intelligence of Open-Source LLMS
topic Computation and Language
url https://arxiv.org/abs/2505.03059