Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Qiyuan, Huang, Hongsen, Shao, Qian, Chen, Jiahe, Chen, Jintai, Xu, Hongxia, Hua, Renjie, Chuan, Ren, Wu, Jian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911141097111552
author Chen, Qiyuan
Huang, Hongsen
Shao, Qian
Chen, Jiahe
Chen, Jintai
Xu, Hongxia
Hua, Renjie
Chuan, Ren
Wu, Jian
author_facet Chen, Qiyuan
Huang, Hongsen
Shao, Qian
Chen, Jiahe
Chen, Jintai
Xu, Hongxia
Hua, Renjie
Chuan, Ren
Wu, Jian
contents Large Language Models (LLMs) require high quality preference datasets to align with human preferences. However, conventional methods for constructing such datasets face significant challenges: reliance on pre-collected instructions often leads to distribution mismatches with target models, while the need for sampling multiple stochastic responses introduces substantial computational overhead. In this work, we explore a paradigm shift by leveraging inherent regulation of LLMs' representation space for efficient and tailored preference dataset construction, named Icon$^{2}$. Specifically, it first extracts layer-wise direction vectors to encode sophisticated human preferences and then uses these vectors to filter self-synthesized instructions based on their inherent consistency. During decoding, bidirectional inherent control is applied to steer token representations, enabling the precise generation of response pairs with clear alignment distinctions. Experimental results demonstrate significant improvements in both alignment and efficiency. Llama3-8B and Qwen2-7B achieve an average win rate improvement of 13.89% on AlpacaEval 2.0 and 13.45% on Arena-Hard, while reducing computational costs by up to 48.1%.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05605
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation
Chen, Qiyuan
Huang, Hongsen
Shao, Qian
Chen, Jiahe
Chen, Jintai
Xu, Hongxia
Hua, Renjie
Chuan, Ren
Wu, Jian
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) require high quality preference datasets to align with human preferences. However, conventional methods for constructing such datasets face significant challenges: reliance on pre-collected instructions often leads to distribution mismatches with target models, while the need for sampling multiple stochastic responses introduces substantial computational overhead. In this work, we explore a paradigm shift by leveraging inherent regulation of LLMs' representation space for efficient and tailored preference dataset construction, named Icon$^{2}$. Specifically, it first extracts layer-wise direction vectors to encode sophisticated human preferences and then uses these vectors to filter self-synthesized instructions based on their inherent consistency. During decoding, bidirectional inherent control is applied to steer token representations, enabling the precise generation of response pairs with clear alignment distinctions. Experimental results demonstrate significant improvements in both alignment and efficiency. Llama3-8B and Qwen2-7B achieve an average win rate improvement of 13.89% on AlpacaEval 2.0 and 13.45% on Arena-Hard, while reducing computational costs by up to 48.1%.
title Icon$^{2}$: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.05605