Alignment Data Map for Efficient Preference Data Selection and Diagnosis

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lee, Seohyeong, Kim, Eunwon, Lee, Hwaran, Chang, Buru
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911606580969472
author Lee, Seohyeong
Kim, Eunwon
Lee, Hwaran
Chang, Buru
author_facet Lee, Seohyeong
Kim, Eunwon
Lee, Hwaran
Chang, Buru
contents Human preference data is essential for aligning large language models (LLMs) with human values, but collecting such data is often costly and inefficient-motivating the need for efficient data selection methods that reduce annotation costs while preserving alignment effectiveness. To address this issue, we propose Alignment Data Map, a data analysis tool for identifying and selecting effective preference data. We first evaluate alignment scores of the preference data by LLM-as-a-judge, explicit reward model, and reference-based approaches. The Alignment Data Map considers both response quality and inter-response variability based on the alignment scores. From our experimental findings, training on only 33% of samples that exhibit high-quality and low-variability, achieves comparable or superior alignment performance on MT-Bench, Evol-Instruct, and AlpacaEval, compared to training with the full dataset. In addition, Alignment Data Map detects potential label misannotations by analyzing correlations between annotated labels and alignment scores, improving annotation accuracy. The implementation is available at https://github.com/01choco/Alignment-Data-Map.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23114
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Alignment Data Map for Efficient Preference Data Selection and Diagnosis
Lee, Seohyeong
Kim, Eunwon
Lee, Hwaran
Chang, Buru
Computation and Language
Human preference data is essential for aligning large language models (LLMs) with human values, but collecting such data is often costly and inefficient-motivating the need for efficient data selection methods that reduce annotation costs while preserving alignment effectiveness. To address this issue, we propose Alignment Data Map, a data analysis tool for identifying and selecting effective preference data. We first evaluate alignment scores of the preference data by LLM-as-a-judge, explicit reward model, and reference-based approaches. The Alignment Data Map considers both response quality and inter-response variability based on the alignment scores. From our experimental findings, training on only 33% of samples that exhibit high-quality and low-variability, achieves comparable or superior alignment performance on MT-Bench, Evol-Instruct, and AlpacaEval, compared to training with the full dataset. In addition, Alignment Data Map detects potential label misannotations by analyzing correlations between annotated labels and alignment scores, improving annotation accuracy. The implementation is available at https://github.com/01choco/Alignment-Data-Map.
title Alignment Data Map for Efficient Preference Data Selection and Diagnosis
topic Computation and Language
url https://arxiv.org/abs/2505.23114