ToolRM: Outcome Reward Models for Tool-Calling Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Agarwal, Mayank, Abdelaziz, Ibrahim, Basu, Kinjal, Unuvar, Merve, Lastras, Luis A., Rizk, Yara, Kapanipathi, Pavan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915713649737728
author Agarwal, Mayank
Abdelaziz, Ibrahim
Basu, Kinjal
Unuvar, Merve
Lastras, Luis A.
Rizk, Yara
Kapanipathi, Pavan
author_facet Agarwal, Mayank
Abdelaziz, Ibrahim
Basu, Kinjal
Unuvar, Merve
Lastras, Luis A.
Rizk, Yara
Kapanipathi, Pavan
contents As large language models (LLMs) increasingly interact with external tools, reward modeling for tool use has emerged as a critical yet underexplored area of research. Existing reward models, trained primarily on natural language outputs, struggle to evaluate tool-based reasoning and execution. To quantify this gap, we introduce FC-RewardBench, the first benchmark to systematically evaluate reward models in tool-calling scenarios. Our analysis shows that current reward models frequently miss key signals of effective tool use, highlighting the need for domain-specific modeling. We address this by proposing a training framework for outcome reward models using data synthesized from permissively licensed, open-weight LLMs. We introduce ToolRM - a suite of reward models for tool-use ranging from 1.7B to 14B parameters. Across diverse settings, these models consistently outperform general-purpose baselines. Notably, they achieve up to a 25% improvement with Best-of-N sampling, while also improving robustness to input noise, enabling effective data filtering, and supporting RL-training of policy models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11963
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
Agarwal, Mayank
Abdelaziz, Ibrahim
Basu, Kinjal
Unuvar, Merve
Lastras, Luis A.
Rizk, Yara
Kapanipathi, Pavan
Computation and Language
As large language models (LLMs) increasingly interact with external tools, reward modeling for tool use has emerged as a critical yet underexplored area of research. Existing reward models, trained primarily on natural language outputs, struggle to evaluate tool-based reasoning and execution. To quantify this gap, we introduce FC-RewardBench, the first benchmark to systematically evaluate reward models in tool-calling scenarios. Our analysis shows that current reward models frequently miss key signals of effective tool use, highlighting the need for domain-specific modeling. We address this by proposing a training framework for outcome reward models using data synthesized from permissively licensed, open-weight LLMs. We introduce ToolRM - a suite of reward models for tool-use ranging from 1.7B to 14B parameters. Across diverse settings, these models consistently outperform general-purpose baselines. Notably, they achieve up to a 25% improvement with Best-of-N sampling, while also improving robustness to input noise, enabling effective data filtering, and supporting RL-training of policy models.
title ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2509.11963