AIR: A Systematic Analysis of Annotations, Instructions, and Response Pairs in Preference Dataset

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: He, Bingxiang, Zhang, Wenbin, Song, Jiaxi, Qian, Cheng, Fu, Zixuan, Sun, Bowen, Ding, Ning, Hong, Haiwen, Huang, Longtao, Xue, Hui, Cui, Ganqu, Che, Wanxiang, Liu, Zhiyuan, Sun, Maosong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918133252489216
author He, Bingxiang
Zhang, Wenbin
Song, Jiaxi
Qian, Cheng
Fu, Zixuan
Sun, Bowen
Ding, Ning
Hong, Haiwen
Huang, Longtao
Xue, Hui
Cui, Ganqu
Che, Wanxiang
Liu, Zhiyuan
Sun, Maosong
author_facet He, Bingxiang
Zhang, Wenbin
Song, Jiaxi
Qian, Cheng
Fu, Zixuan
Sun, Bowen
Ding, Ning
Hong, Haiwen
Huang, Longtao
Xue, Hui
Cui, Ganqu
Che, Wanxiang
Liu, Zhiyuan
Sun, Maosong
contents Preference learning is critical for aligning large language models (LLMs) with human values, yet its success hinges on high-quality datasets comprising three core components: Preference \textbf{A}nnotations, \textbf{I}nstructions, and \textbf{R}esponse Pairs. Current approaches conflate these components, obscuring their individual impacts and hindering systematic optimization. In this work, we propose \textbf{AIR}, a component-wise analysis framework that systematically isolates and optimizes each component while evaluating their synergistic effects. Through rigorous experimentation, AIR reveals actionable principles: annotation simplicity (point-wise generative scoring), instruction inference stability (variance-based filtering across LLMs), and response pair quality (moderate margins + high absolute scores). When combined, these principles yield +5.3 average gains over baseline method, even with only 14k high-quality pairs. Our work shifts preference dataset design from ad hoc scaling to component-aware optimization, offering a blueprint for efficient, reproducible alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2504_03612
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AIR: A Systematic Analysis of Annotations, Instructions, and Response Pairs in Preference Dataset
He, Bingxiang
Zhang, Wenbin
Song, Jiaxi
Qian, Cheng
Fu, Zixuan
Sun, Bowen
Ding, Ning
Hong, Haiwen
Huang, Longtao
Xue, Hui
Cui, Ganqu
Che, Wanxiang
Liu, Zhiyuan
Sun, Maosong
Computation and Language
Preference learning is critical for aligning large language models (LLMs) with human values, yet its success hinges on high-quality datasets comprising three core components: Preference \textbf{A}nnotations, \textbf{I}nstructions, and \textbf{R}esponse Pairs. Current approaches conflate these components, obscuring their individual impacts and hindering systematic optimization. In this work, we propose \textbf{AIR}, a component-wise analysis framework that systematically isolates and optimizes each component while evaluating their synergistic effects. Through rigorous experimentation, AIR reveals actionable principles: annotation simplicity (point-wise generative scoring), instruction inference stability (variance-based filtering across LLMs), and response pair quality (moderate margins + high absolute scores). When combined, these principles yield +5.3 average gains over baseline method, even with only 14k high-quality pairs. Our work shifts preference dataset design from ad hoc scaling to component-aware optimization, offering a blueprint for efficient, reproducible alignment.
title AIR: A Systematic Analysis of Annotations, Instructions, and Response Pairs in Preference Dataset
topic Computation and Language
url https://arxiv.org/abs/2504.03612