Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Shi, Wang, Yingjie, Hu, Shengchao, Wang, Peng, Tao, Dacheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912870464225280
author Fu, Shi
Wang, Yingjie
Hu, Shengchao
Wang, Peng
Tao, Dacheng
author_facet Fu, Shi
Wang, Yingjie
Hu, Shengchao
Wang, Peng
Tao, Dacheng
contents Self-Rewarding Language Models (SRLMs) achieve notable success in iteratively improving alignment without external feedback. Yet, despite their striking empirical progress, the core mechanisms driving their capabilities remain unelucidated, leaving a critical gap in theoretical understanding. This paper provides the first rigorous theoretical guarantees for SRLMs. We first establish a lower bound that characterizes the fundamental limits of a single update step, revealing a critical dependence on the quality of the initial model. We then derive finite-sample error bounds for the full iterative paradigm, showing that performance improves at a rate of $\widetilde{\mathcal{O}}\left(1/\sqrt{n}\right)$ with sample size $n$. Crucially, our analysis reveals that the dependence on the initial model decays exponentially with the number of iterations $T$. This provides a formal explanation for why self-rewarding succeeds: it robustly overcomes poor initialization by steering the dynamics toward internal stability and consistency. Finally, we instantiate our theoretical framework for the linear softmax model class, yielding tailored guarantees that connect our high-level insights to practical model architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22513
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language Models
Fu, Shi
Wang, Yingjie
Hu, Shengchao
Wang, Peng
Tao, Dacheng
Artificial Intelligence
Self-Rewarding Language Models (SRLMs) achieve notable success in iteratively improving alignment without external feedback. Yet, despite their striking empirical progress, the core mechanisms driving their capabilities remain unelucidated, leaving a critical gap in theoretical understanding. This paper provides the first rigorous theoretical guarantees for SRLMs. We first establish a lower bound that characterizes the fundamental limits of a single update step, revealing a critical dependence on the quality of the initial model. We then derive finite-sample error bounds for the full iterative paradigm, showing that performance improves at a rate of $\widetilde{\mathcal{O}}\left(1/\sqrt{n}\right)$ with sample size $n$. Crucially, our analysis reveals that the dependence on the initial model decays exponentially with the number of iterations $T$. This provides a formal explanation for why self-rewarding succeeds: it robustly overcomes poor initialization by steering the dynamics toward internal stability and consistency. Finally, we instantiate our theoretical framework for the linear softmax model class, yielding tailored guarantees that connect our high-level insights to practical model architectures.
title Why Self-Rewarding Works: Theoretical Guarantees for Iterative Alignment of Language Models
topic Artificial Intelligence
url https://arxiv.org/abs/2601.22513