SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Yuru, Ding, Wenxuan, Feng, Shangbin, Durrett, Greg, Tsvetkov, Yulia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908623686336512
author Jiang, Yuru
Ding, Wenxuan
Feng, Shangbin
Durrett, Greg
Tsvetkov, Yulia
author_facet Jiang, Yuru
Ding, Wenxuan
Feng, Shangbin
Durrett, Greg
Tsvetkov, Yulia
contents We propose SPARTA ALIGNMENT, an algorithm to collectively align multiple LLMs through competition and combat. To complement a single model's lack of diversity in generation and biases in evaluation, multiple LLMs form a "sparta tribe" to compete against each other in fulfilling instructions while serving as judges for the competition of others. For each iteration, one instruction and two models are selected for a duel, the other models evaluate the two responses, and their evaluation scores are aggregated through a adapted elo-ranking based reputation system, where winners/losers of combat gain/lose weight in evaluating others. The peer-evaluated combat results then become preference pairs where the winning response is preferred over the losing one, and all models learn from these preferences at the end of each iteration. SPARTA ALIGNMENT enables the self-evolution of multiple LLMs in an iterative and collective competition process. Extensive experiments demonstrate that SPARTA ALIGNMENT outperforms initial models and 4 self-alignment baselines across 10 out of 12 tasks and datasets with 7.0% average improvement. Further analysis reveals that SPARTA ALIGNMENT generalizes more effectively to unseen tasks and leverages the expertise diversity of participating models to produce more logical, direct and informative outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04721
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat
Jiang, Yuru
Ding, Wenxuan
Feng, Shangbin
Durrett, Greg
Tsvetkov, Yulia
Computation and Language
We propose SPARTA ALIGNMENT, an algorithm to collectively align multiple LLMs through competition and combat. To complement a single model's lack of diversity in generation and biases in evaluation, multiple LLMs form a "sparta tribe" to compete against each other in fulfilling instructions while serving as judges for the competition of others. For each iteration, one instruction and two models are selected for a duel, the other models evaluate the two responses, and their evaluation scores are aggregated through a adapted elo-ranking based reputation system, where winners/losers of combat gain/lose weight in evaluating others. The peer-evaluated combat results then become preference pairs where the winning response is preferred over the losing one, and all models learn from these preferences at the end of each iteration. SPARTA ALIGNMENT enables the self-evolution of multiple LLMs in an iterative and collective competition process. Extensive experiments demonstrate that SPARTA ALIGNMENT outperforms initial models and 4 self-alignment baselines across 10 out of 12 tasks and datasets with 7.0% average improvement. Further analysis reveals that SPARTA ALIGNMENT generalizes more effectively to unseen tasks and leverages the expertise diversity of participating models to produce more logical, direct and informative outputs.
title SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat
topic Computation and Language
url https://arxiv.org/abs/2506.04721