TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Dan, Cai, Min, Light, Jonathan, Hu, Ziniu, Yue, Yisong, Tang, Jie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908564894777344
author Zhang, Dan
Cai, Min
Light, Jonathan
Hu, Ziniu
Yue, Yisong
Tang, Jie
author_facet Zhang, Dan
Cai, Min
Light, Jonathan
Hu, Ziniu
Yue, Yisong
Tang, Jie
contents Reward models are central to both reinforcement learning (RL) with language models and inference-time verification. However, existing reward models often lack temporal consistency, leading to ineffective policy updates and unstable RL training. We introduce TDRM, a method for learning smoother and more reliable reward models by minimizing temporal differences (TD) for training-time reinforcement learning and inference-time verification. Experiments show that TD-trained process reward models (PRMs) improve performance across Best-of-N (up to 6.6%) and tree-search (up to 23.7%) settings. When combined with Reinforcement Learning with Verifiable Rewards (RLVR), TD-trained PRMs lead to more data-efficient RL -- achieving comparable performance with just 2.5k data to what baseline methods require 50.1k data to attain -- and yield higher-quality language model policies in 8 model variants (5 series), e.g., Qwen2.5-(0.5B, 1,5B), GLM4-9B-0414, GLM-Z1-9B-0414, Qwen2.5-Math-(1.5B, 7B), and DeepSeek-R1-Distill-Qwen-(1.5B, 7B). We release all code at https://github.com/THUDM/TDRM.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15110
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
Zhang, Dan
Cai, Min
Light, Jonathan
Hu, Ziniu
Yue, Yisong
Tang, Jie
Machine Learning
Computation and Language
Reward models are central to both reinforcement learning (RL) with language models and inference-time verification. However, existing reward models often lack temporal consistency, leading to ineffective policy updates and unstable RL training. We introduce TDRM, a method for learning smoother and more reliable reward models by minimizing temporal differences (TD) for training-time reinforcement learning and inference-time verification. Experiments show that TD-trained process reward models (PRMs) improve performance across Best-of-N (up to 6.6%) and tree-search (up to 23.7%) settings. When combined with Reinforcement Learning with Verifiable Rewards (RLVR), TD-trained PRMs lead to more data-efficient RL -- achieving comparable performance with just 2.5k data to what baseline methods require 50.1k data to attain -- and yield higher-quality language model policies in 8 model variants (5 series), e.g., Qwen2.5-(0.5B, 1,5B), GLM4-9B-0414, GLM-Z1-9B-0414, Qwen2.5-Math-(1.5B, 7B), and DeepSeek-R1-Distill-Qwen-(1.5B, 7B). We release all code at https://github.com/THUDM/TDRM.
title TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2509.15110