MUSIC: MUlti-Step Instruction Contrast for Multi-Turn Reward Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Wenzhe, Zhang, Shujian, Zhou, Wenxuan, Lambert, John, Jin, Chi, Hard, Andrew, Mathews, Rajiv, Wang, Lun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917177957810176
author Li, Wenzhe
Zhang, Shujian
Zhou, Wenxuan
Lambert, John
Jin, Chi
Hard, Andrew
Mathews, Rajiv
Wang, Lun
author_facet Li, Wenzhe
Zhang, Shujian
Zhou, Wenxuan
Lambert, John
Jin, Chi
Hard, Andrew
Mathews, Rajiv
Wang, Lun
contents Evaluating the quality of multi-turn conversations is crucial for developing capable Large Language Models (LLMs), yet remains a significant challenge, often requiring costly human evaluation. Multi-turn reward models (RMs) offer a scalable alternative and can provide valuable signals for guiding LLM training. While recent work has advanced multi-turn \textit{training} techniques, effective automated \textit{evaluation} specifically for multi-turn interactions lags behind. We observe that standard preference datasets, typically contrasting responses based only on the final conversational turn, provide insufficient signal to capture the nuances of multi-turn interactions. Instead, we find that incorporating contrasts spanning \textit{multiple} turns is critical for building robust multi-turn RMs. Motivated by this finding, we propose \textbf{MU}lti-\textbf{S}tep \textbf{I}nstruction \textbf{C}ontrast (MUSIC), an unsupervised data augmentation strategy that synthesizes contrastive conversation pairs exhibiting differences across multiple turns. Leveraging MUSIC on the Skywork preference dataset, we train a multi-turn RM based on the Gemma-2-9B-Instruct model. Empirical results demonstrate that our MUSIC-augmented RM outperforms baseline methods, achieving higher alignment with judgments from advanced proprietary LLM judges on multi-turn conversations, crucially, without compromising performance on standard single-turn RM benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24693
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MUSIC: MUlti-Step Instruction Contrast for Multi-Turn Reward Models
Li, Wenzhe
Zhang, Shujian
Zhou, Wenxuan
Lambert, John
Jin, Chi
Hard, Andrew
Mathews, Rajiv
Wang, Lun
Computation and Language
Evaluating the quality of multi-turn conversations is crucial for developing capable Large Language Models (LLMs), yet remains a significant challenge, often requiring costly human evaluation. Multi-turn reward models (RMs) offer a scalable alternative and can provide valuable signals for guiding LLM training. While recent work has advanced multi-turn \textit{training} techniques, effective automated \textit{evaluation} specifically for multi-turn interactions lags behind. We observe that standard preference datasets, typically contrasting responses based only on the final conversational turn, provide insufficient signal to capture the nuances of multi-turn interactions. Instead, we find that incorporating contrasts spanning \textit{multiple} turns is critical for building robust multi-turn RMs. Motivated by this finding, we propose \textbf{MU}lti-\textbf{S}tep \textbf{I}nstruction \textbf{C}ontrast (MUSIC), an unsupervised data augmentation strategy that synthesizes contrastive conversation pairs exhibiting differences across multiple turns. Leveraging MUSIC on the Skywork preference dataset, we train a multi-turn RM based on the Gemma-2-9B-Instruct model. Empirical results demonstrate that our MUSIC-augmented RM outperforms baseline methods, achieving higher alignment with judgments from advanced proprietary LLM judges on multi-turn conversations, crucially, without compromising performance on standard single-turn RM benchmarks.
title MUSIC: MUlti-Step Instruction Contrast for Multi-Turn Reward Models
topic Computation and Language
url https://arxiv.org/abs/2512.24693