Diagnosing Training Inference Mismatch in LLM Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911683918692352 |
|---|---|
| author | Zhong, Tianle Ling, Neiwen Pi, Yifan Wei, Zijun Yu, Tianshu Fox, Geoffrey Wu, Peng Yu, Xiao |
| author_facet | Zhong, Tianle Ling, Neiwen Pi, Yifan Wei, Zijun Yu, Tianshu Fox, Geoffrey Wu, Peng Yu, Xiao |
| contents | Modern LLM RL systems separate rollout generation from policy optimization. These two stages are expected to produce token probabilities that match exactly. However, implementation differences can make them assign different values to the same sequence under the same model weights, inducing Training-Inference Mismatch (TIM). TIM is difficult to inspect because it is entangled with off-policy drift and common stabilization mechanisms. In this work, we isolate TIM in a zero-mismatch diagnostic setting (VeXact), and show that small token-level numerical disagreements can independently cause training collapse. We further show that TIM changes the effective optimization problem, and identify a set of remedies that could mitigate TIM. Our results suggest that TIM is not benign numerical noise, but a systems-level perturbation that should be treated as a first-order factor in analyzing LLM RL stability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_14220 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Diagnosing Training Inference Mismatch in LLM Reinforcement Learning Zhong, Tianle Ling, Neiwen Pi, Yifan Wei, Zijun Yu, Tianshu Fox, Geoffrey Wu, Peng Yu, Xiao Machine Learning Artificial Intelligence Computation and Language Modern LLM RL systems separate rollout generation from policy optimization. These two stages are expected to produce token probabilities that match exactly. However, implementation differences can make them assign different values to the same sequence under the same model weights, inducing Training-Inference Mismatch (TIM). TIM is difficult to inspect because it is entangled with off-policy drift and common stabilization mechanisms. In this work, we isolate TIM in a zero-mismatch diagnostic setting (VeXact), and show that small token-level numerical disagreements can independently cause training collapse. We further show that TIM changes the effective optimization problem, and identify a set of remedies that could mitigate TIM. Our results suggest that TIM is not benign numerical noise, but a systems-level perturbation that should be treated as a first-order factor in analyzing LLM RL stability. |
| title | Diagnosing Training Inference Mismatch in LLM Reinforcement Learning |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2605.14220 |