Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhou, Xin, Guo, Yiwen, Ma, Ruotian, Gui, Tao, Zhang, Qi, Huang, Xuanjing
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912230910459904
author Zhou, Xin
Guo, Yiwen
Ma, Ruotian
Gui, Tao
Zhang, Qi
Huang, Xuanjing
author_facet Zhou, Xin
Guo, Yiwen
Ma, Ruotian
Gui, Tao
Zhang, Qi
Huang, Xuanjing
contents Aligning Large Language Models (LLMs) with human preferences is crucial for their deployment in real-world applications. Recent advancements in Self-Rewarding Language Models suggest that an LLM can use its internal reward models (such as LLM-as-a-Judge) \cite{yuanself} to generate preference data, improving alignment performance without costly human annotation. However, we find that different internal reward models within the same LLM often generate inconsistent preferences. This inconsistency raises concerns about the reliability of self-generated preference data, hinders overall alignment performance, and highlights the need for further research to ensure reliable and coherent alignment with human preferences. To address this limitation, we propose Self-Consistent Internal Rewards (SCIR), a novel framework designed to enhance consistency among internal reward models during training. In each training step, we collect preference predictions from multiple pre-defined internal reward models and enforce consistency and confidence through an inconsistency penalty mechanism, thereby improving the reliability of these internal reward models. We selectively use data with consistent predictions for preference optimization, ensuring the quality of the preference data. By employing self-consistent internal rewards, our method significantly improves the alignment performance and reward modeling capability of LLMs, outperforming baseline methods by a notable margin.
format Preprint
id arxiv_https___arxiv_org_abs_2502_08922
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
Zhou, Xin
Guo, Yiwen
Ma, Ruotian
Gui, Tao
Zhang, Qi
Huang, Xuanjing
Artificial Intelligence
Aligning Large Language Models (LLMs) with human preferences is crucial for their deployment in real-world applications. Recent advancements in Self-Rewarding Language Models suggest that an LLM can use its internal reward models (such as LLM-as-a-Judge) \cite{yuanself} to generate preference data, improving alignment performance without costly human annotation. However, we find that different internal reward models within the same LLM often generate inconsistent preferences. This inconsistency raises concerns about the reliability of self-generated preference data, hinders overall alignment performance, and highlights the need for further research to ensure reliable and coherent alignment with human preferences. To address this limitation, we propose Self-Consistent Internal Rewards (SCIR), a novel framework designed to enhance consistency among internal reward models during training. In each training step, we collect preference predictions from multiple pre-defined internal reward models and enforce consistency and confidence through an inconsistency penalty mechanism, thereby improving the reliability of these internal reward models. We selectively use data with consistent predictions for preference optimization, ensuring the quality of the preference data. By employing self-consistent internal rewards, our method significantly improves the alignment performance and reward modeling capability of LLMs, outperforming baseline methods by a notable margin.
title Self-Consistency of the Internal Reward Models Improves Self-Rewarding Language Models
topic Artificial Intelligence
url https://arxiv.org/abs/2502.08922