Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Fei, Zhang, Yongkang, Liu, Runhao, Liao, Yuhao, Zeng, Zijian, Yang, Huiming, wang, Sibo, Liao, Linglin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913159064846336
author Ding, Fei
Zhang, Yongkang
Liu, Runhao
Liao, Yuhao
Zeng, Zijian
Yang, Huiming
wang, Sibo
Liao, Linglin
author_facet Ding, Fei
Zhang, Yongkang
Liu, Runhao
Liao, Yuhao
Zeng, Zijian
Yang, Huiming
wang, Sibo
Liao, Linglin
contents This paper investigates the length problem in sequence-level relative reinforcement learning. We observe that, although existing methods partially alleviate length-related phenomena, a more fundamental issue remains insufficiently characterized: the comparison units used during training lack inherent comparability. Building on this observation, we propose a new perspective: the length problem should not be viewed merely as a loss-scaling or normalization bias, but rather as a \emph{comparison unit construction} problem. We further establish a sample-construction-based training framework that, instead of applying post-hoc corrections to unequal-length responses, proactively constructs equal-length, alignable, and comparable training segments during generation. Within this framework, we propose EqLen, a concrete method applicable to group-relative comparison algorithms such as GRPO, GSPO, and RLOO. Through dual-track synchronous generation, prefix inheritance, and segment masking, EqLen efficiently collects effective equal-length training segments and enables stable
format Preprint
id arxiv_https___arxiv_org_abs_2604_17328
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction
Ding, Fei
Zhang, Yongkang
Liu, Runhao
Liao, Yuhao
Zeng, Zijian
Yang, Huiming
wang, Sibo
Liao, Linglin
Machine Learning
Artificial Intelligence
This paper investigates the length problem in sequence-level relative reinforcement learning. We observe that, although existing methods partially alleviate length-related phenomena, a more fundamental issue remains insufficiently characterized: the comparison units used during training lack inherent comparability. Building on this observation, we propose a new perspective: the length problem should not be viewed merely as a loss-scaling or normalization bias, but rather as a \emph{comparison unit construction} problem. We further establish a sample-construction-based training framework that, instead of applying post-hoc corrections to unequal-length responses, proactively constructs equal-length, alignable, and comparable training segments during generation. Within this framework, we propose EqLen, a concrete method applicable to group-relative comparison algorithms such as GRPO, GSPO, and RLOO. Through dual-track synchronous generation, prefix inheritance, and segment masking, EqLen efficiently collects effective equal-length training segments and enables stable
title Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2604.17328