Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Ran, Chen, Jingjing, Ye, Jiayu, Wu, Yu, Yan, Jun, Yang, Carl, Yu, Hongkun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915810611560448
author Xu, Ran
Chen, Jingjing
Ye, Jiayu
Wu, Yu
Yan, Jun
Yang, Carl
Yu, Hongkun
author_facet Xu, Ran
Chen, Jingjing
Ye, Jiayu
Wu, Yu
Yan, Jun
Yang, Carl
Yu, Hongkun
contents Large Language Models (LLMs) are widely used as judges to evaluate response quality, providing a scalable alternative to human evaluation. However, most LLM judges operate solely on intrinsic text-based reasoning, limiting their ability to verify complex constraints or perform accurate computation. Motivated by the success of tool-integrated reasoning (TIR) in numerous tasks, we propose TIR-Judge, an end-to-end RL framework for training LLM judges that integrates a code executor for precise evaluation. TIR-Judge is built on three principles: (i) diverse training across verifiable and non-verifiable domains, (ii) flexible judgment formats (pointwise, pairwise, listwise), and (iii) iterative RL that bootstraps directly from the initial model without distillation. On seven public benchmarks, TIR-Judge surpasses strong reasoning-based judges by up to 6.4% (pointwise) and 7.7% (pairwise), and achieves listwise performance comparable to Claude-Opus-4 despite having only 8B parameters. Remarkably, TIR-Judge-Zero - trained entirely without distilled judge trajectories, matches the performance of distilled variants, demonstrating that tool-augmented judges can self-evolve through iterative reinforcement learning.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23038
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
Xu, Ran
Chen, Jingjing
Ye, Jiayu
Wu, Yu
Yan, Jun
Yang, Carl
Yu, Hongkun
Computation and Language
Artificial Intelligence
Machine Learning
Large Language Models (LLMs) are widely used as judges to evaluate response quality, providing a scalable alternative to human evaluation. However, most LLM judges operate solely on intrinsic text-based reasoning, limiting their ability to verify complex constraints or perform accurate computation. Motivated by the success of tool-integrated reasoning (TIR) in numerous tasks, we propose TIR-Judge, an end-to-end RL framework for training LLM judges that integrates a code executor for precise evaluation. TIR-Judge is built on three principles: (i) diverse training across verifiable and non-verifiable domains, (ii) flexible judgment formats (pointwise, pairwise, listwise), and (iii) iterative RL that bootstraps directly from the initial model without distillation. On seven public benchmarks, TIR-Judge surpasses strong reasoning-based judges by up to 6.4% (pointwise) and 7.7% (pairwise), and achieves listwise performance comparable to Claude-Opus-4 despite having only 8B parameters. Remarkably, TIR-Judge-Zero - trained entirely without distilled judge trajectories, matches the performance of distilled variants, demonstrating that tool-augmented judges can self-evolve through iterative reinforcement learning.
title Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.23038