Refine-n-Judge: Curating High-Quality Preference Chains for LLM-Fine-Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cayir, Derin, Tao, Renjie, Rungta, Rashi, Sun, Kai, Chen, Sean, Khan, Haidar, Kim, Minseok, Reinspach, Julia, Liu, Yue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911240045985792
author Cayir, Derin
Tao, Renjie
Rungta, Rashi
Sun, Kai
Chen, Sean
Khan, Haidar
Kim, Minseok
Reinspach, Julia
Liu, Yue
author_facet Cayir, Derin
Tao, Renjie
Rungta, Rashi
Sun, Kai
Chen, Sean
Khan, Haidar
Kim, Minseok
Reinspach, Julia
Liu, Yue
contents Large Language Models (LLMs) have demonstrated remarkable progress through preference-based fine-tuning, which critically depends on the quality of the underlying training data. While human feedback is essential for improving data quality, it is costly and does not scale well. In this paper, we introduce Refine-n-Judge, an automated iterative approach that leverages a single LLM as both a refiner and a judge to enhance dataset quality. Unlike existing iterative refinement methods, Refine-n-Judge employs an LLM to both generate refinements and explicitly evaluate each improvement, ensuring that every iteration meaningfully enhances the dataset without requiring additional human annotation or a separate reward model. At each step, the LLM refines a response and judges whether the refinement is an improvement over the previous answer. This process continues until the LLM prefers the initial answer over the refinement, indicating no further improvements. This produces sequences of increasing quality, preference-labeled responses ideal for fine-tuning. We demonstrate the effectiveness of Refine-n-Judge across a range of public datasets spanning five corpora, targeting tasks such as coding, math, and conversation. Models (Llama 3.1-8B and Llama 3.3-70B) fine-tuned on Refine-n-Judge-enhanced datasets were preferred by LLM judges in over 74% of comparisons against models tuned on the original dataset by GPT-4. Additionally, we report performance gains: +5% on AlpacaEval and AlpacaEval 2.0, and +19% on MT-Bench. Our results indicate that Refine-n-Judge produces high-quality datasets and scalable model improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01543
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Refine-n-Judge: Curating High-Quality Preference Chains for LLM-Fine-Tuning
Cayir, Derin
Tao, Renjie
Rungta, Rashi
Sun, Kai
Chen, Sean
Khan, Haidar
Kim, Minseok
Reinspach, Julia
Liu, Yue
Artificial Intelligence
Large Language Models (LLMs) have demonstrated remarkable progress through preference-based fine-tuning, which critically depends on the quality of the underlying training data. While human feedback is essential for improving data quality, it is costly and does not scale well. In this paper, we introduce Refine-n-Judge, an automated iterative approach that leverages a single LLM as both a refiner and a judge to enhance dataset quality. Unlike existing iterative refinement methods, Refine-n-Judge employs an LLM to both generate refinements and explicitly evaluate each improvement, ensuring that every iteration meaningfully enhances the dataset without requiring additional human annotation or a separate reward model. At each step, the LLM refines a response and judges whether the refinement is an improvement over the previous answer. This process continues until the LLM prefers the initial answer over the refinement, indicating no further improvements. This produces sequences of increasing quality, preference-labeled responses ideal for fine-tuning. We demonstrate the effectiveness of Refine-n-Judge across a range of public datasets spanning five corpora, targeting tasks such as coding, math, and conversation. Models (Llama 3.1-8B and Llama 3.3-70B) fine-tuned on Refine-n-Judge-enhanced datasets were preferred by LLM judges in over 74% of comparisons against models tuned on the original dataset by GPT-4. Additionally, we report performance gains: +5% on AlpacaEval and AlpacaEval 2.0, and +19% on MT-Bench. Our results indicate that Refine-n-Judge produces high-quality datasets and scalable model improvements.
title Refine-n-Judge: Curating High-Quality Preference Chains for LLM-Fine-Tuning
topic Artificial Intelligence
url https://arxiv.org/abs/2508.01543