Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zili, Chai, Jiajun, Chen, Lin, Wang, Xiaohan, Xiang, Shiming, Yin, Guojun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913166971109376
author Wang, Zili
Chai, Jiajun
Chen, Lin
Wang, Xiaohan
Xiang, Shiming
Yin, Guojun
author_facet Wang, Zili
Chai, Jiajun
Chen, Lin
Wang, Xiaohan
Xiang, Shiming
Yin, Guojun
contents Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as the standard paradigm for improving reasoning capability of large language models, while Multi-Token Prediction (MTP) has been a widely adopted module in pretraining. Combining them is a natural approach, yet current RL practices detach MTP gradients because joint training degrades the performance. We revisit this failure from an optimization perspective. We show that the per-step effect of MTP on the RL objective can be decomposed into two terms: a first-order correlation and a second-order perturbation penalty. This decomposition unifies three MTP training regimes: Detach, Cross-Entropy loss, and Policy loss, and explains why each succeeds or fails. Further analysis of policy loss reveals that, although it aligns with intuition, performance still degrades: the correlation term decays while the quadratic penalty persists. Guided by the analysis, we propose Optimal Coefficient Calibration (OCC), an adaptive scheme that tracks the optimal coefficient online via a log-probability proxy at negligible cost. Across six competition-level mathematical reasoning benchmarks, OCC consistently matches or exceeds the detach baseline, delivering improved joint MTP-RL training performance.
format Preprint
id arxiv_https___arxiv_org_abs_2605_28184
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration
Wang, Zili
Chai, Jiajun
Chen, Lin
Wang, Xiaohan
Xiang, Shiming
Yin, Guojun
Machine Learning
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as the standard paradigm for improving reasoning capability of large language models, while Multi-Token Prediction (MTP) has been a widely adopted module in pretraining. Combining them is a natural approach, yet current RL practices detach MTP gradients because joint training degrades the performance. We revisit this failure from an optimization perspective. We show that the per-step effect of MTP on the RL objective can be decomposed into two terms: a first-order correlation and a second-order perturbation penalty. This decomposition unifies three MTP training regimes: Detach, Cross-Entropy loss, and Policy loss, and explains why each succeeds or fails. Further analysis of policy loss reveals that, although it aligns with intuition, performance still degrades: the correlation term decays while the quadratic penalty persists. Guided by the analysis, we propose Optimal Coefficient Calibration (OCC), an adaptive scheme that tracks the optimal coefficient online via a log-probability proxy at negligible cost. Across six competition-level mathematical reasoning benchmarks, OCC consistently matches or exceeds the detach baseline, delivering improved joint MTP-RL training performance.
title Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration
topic Machine Learning
url https://arxiv.org/abs/2605.28184