Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Chen, Ma, Yiyuan, Yang, Yuan, Liu, Deyi, Liu, Jing, Song, Zuquan, Song, Yuxin, Ren, Cheng, Zhu, Hang, Liu, Xin, Qiao, Siyuan, Zhou, Xun, Xiang, Liang, Wu, Yonghui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915470834139136
author Zheng, Chen
Ma, Yiyuan
Yang, Yuan
Liu, Deyi
Liu, Jing
Song, Zuquan
Song, Yuxin
Ren, Cheng
Zhu, Hang
Liu, Xin
Ma, Yiyuan
Qiao, Siyuan
Zhou, Xun
Xiang, Liang
Wu, Yonghui
author_facet Zheng, Chen
Ma, Yiyuan
Yang, Yuan
Liu, Deyi
Liu, Jing
Song, Zuquan
Song, Yuxin
Ren, Cheng
Zhu, Hang
Liu, Xin
Ma, Yiyuan
Qiao, Siyuan
Zhou, Xun
Xiang, Liang
Wu, Yonghui
contents The development of alignment and reasoning capabilities in large language models has seen remarkable progress through two paradigms: instruction tuning and reinforcement learning from human feedback (RLHF) alignment paradigm, and distillation-based reasoning fine-tuning paradigm. While both approaches prove effective independently, the third paradigm of applying RLHF to distillation-trained models presents significant challenges. Our investigation reveals two critical phenomena that emerge in this paradigm: Sequence Length Collapse, where language generation dramatically reduces during early RLHF training, and the Reward Hockey Stick Curve, featuring severe reward score drops followed by gradual recovery. These instabilities fundamentally compromise the model's alignment and reasoning capabilities. To address these challenges, we propose Balanced Actor Initialization (BAI), a two-stage weighted model merging approach. BAI first merges instruction-following and distillation-based reasoning fine-tuned models, then further combines this intermediate model with the pretrained model to preserve foundational knowledge. Through comprehensive experiments across diverse benchmarks and detailed analysis of training experiments, we demonstrate that BAI resolves Sequence Length Collapse, mitigates the Reward Hockey Stick Curve, and enables continuous sequence length improvement during training. Additionally, our analysis reveals that balanced merging ratios achieve optimal trade-offs between training stability and reasoning capability preservation. Our work provides the effective solution for stable training in this third paradigm, enabling more capable reasoning models that combine distillation efficiency with RLHF alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00309
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
Zheng, Chen
Ma, Yiyuan
Yang, Yuan
Liu, Deyi
Liu, Jing
Song, Zuquan
Song, Yuxin
Ren, Cheng
Zhu, Hang
Liu, Xin
Ma, Yiyuan
Qiao, Siyuan
Zhou, Xun
Xiang, Liang
Wu, Yonghui
Computation and Language
The development of alignment and reasoning capabilities in large language models has seen remarkable progress through two paradigms: instruction tuning and reinforcement learning from human feedback (RLHF) alignment paradigm, and distillation-based reasoning fine-tuning paradigm. While both approaches prove effective independently, the third paradigm of applying RLHF to distillation-trained models presents significant challenges. Our investigation reveals two critical phenomena that emerge in this paradigm: Sequence Length Collapse, where language generation dramatically reduces during early RLHF training, and the Reward Hockey Stick Curve, featuring severe reward score drops followed by gradual recovery. These instabilities fundamentally compromise the model's alignment and reasoning capabilities. To address these challenges, we propose Balanced Actor Initialization (BAI), a two-stage weighted model merging approach. BAI first merges instruction-following and distillation-based reasoning fine-tuned models, then further combines this intermediate model with the pretrained model to preserve foundational knowledge. Through comprehensive experiments across diverse benchmarks and detailed analysis of training experiments, we demonstrate that BAI resolves Sequence Length Collapse, mitigates the Reward Hockey Stick Curve, and enables continuous sequence length improvement during training. Additionally, our analysis reveals that balanced merging ratios achieve optimal trade-offs between training stability and reasoning capability preservation. Our work provides the effective solution for stable training in this third paradigm, enabling more capable reasoning models that combine distillation efficiency with RLHF alignment.
title Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
topic Computation and Language
url https://arxiv.org/abs/2509.00309