Diagnosing Training Inference Mismatch in LLM Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhong, Tianle, Ling, Neiwen, Pi, Yifan, Wei, Zijun, Yu, Tianshu, Fox, Geoffrey, Wu, Peng, Yu, Xiao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911683918692352
author Zhong, Tianle
Ling, Neiwen
Pi, Yifan
Wei, Zijun
Yu, Tianshu
Fox, Geoffrey
Wu, Peng
Yu, Xiao
author_facet Zhong, Tianle
Ling, Neiwen
Pi, Yifan
Wei, Zijun
Yu, Tianshu
Fox, Geoffrey
Wu, Peng
Yu, Xiao
contents Modern LLM RL systems separate rollout generation from policy optimization. These two stages are expected to produce token probabilities that match exactly. However, implementation differences can make them assign different values to the same sequence under the same model weights, inducing Training-Inference Mismatch (TIM). TIM is difficult to inspect because it is entangled with off-policy drift and common stabilization mechanisms. In this work, we isolate TIM in a zero-mismatch diagnostic setting (VeXact), and show that small token-level numerical disagreements can independently cause training collapse. We further show that TIM changes the effective optimization problem, and identify a set of remedies that could mitigate TIM. Our results suggest that TIM is not benign numerical noise, but a systems-level perturbation that should be treated as a first-order factor in analyzing LLM RL stability.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14220
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Diagnosing Training Inference Mismatch in LLM Reinforcement Learning
Zhong, Tianle
Ling, Neiwen
Pi, Yifan
Wei, Zijun
Yu, Tianshu
Fox, Geoffrey
Wu, Peng
Yu, Xiao
Machine Learning
Artificial Intelligence
Computation and Language
Modern LLM RL systems separate rollout generation from policy optimization. These two stages are expected to produce token probabilities that match exactly. However, implementation differences can make them assign different values to the same sequence under the same model weights, inducing Training-Inference Mismatch (TIM). TIM is difficult to inspect because it is entangled with off-policy drift and common stabilization mechanisms. In this work, we isolate TIM in a zero-mismatch diagnostic setting (VeXact), and show that small token-level numerical disagreements can independently cause training collapse. We further show that TIM changes the effective optimization problem, and identify a set of remedies that could mitigate TIM. Our results suggest that TIM is not benign numerical noise, but a systems-level perturbation that should be treated as a first-order factor in analyzing LLM RL stability.
title Diagnosing Training Inference Mismatch in LLM Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.14220