ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915713649737728 |
|---|---|
| author | Agarwal, Mayank Abdelaziz, Ibrahim Basu, Kinjal Unuvar, Merve Lastras, Luis A. Rizk, Yara Kapanipathi, Pavan |
| author_facet | Agarwal, Mayank Abdelaziz, Ibrahim Basu, Kinjal Unuvar, Merve Lastras, Luis A. Rizk, Yara Kapanipathi, Pavan |
| contents | As large language models (LLMs) increasingly interact with external tools, reward modeling for tool use has emerged as a critical yet underexplored area of research. Existing reward models, trained primarily on natural language outputs, struggle to evaluate tool-based reasoning and execution. To quantify this gap, we introduce FC-RewardBench, the first benchmark to systematically evaluate reward models in tool-calling scenarios. Our analysis shows that current reward models frequently miss key signals of effective tool use, highlighting the need for domain-specific modeling. We address this by proposing a training framework for outcome reward models using data synthesized from permissively licensed, open-weight LLMs. We introduce ToolRM - a suite of reward models for tool-use ranging from 1.7B to 14B parameters. Across diverse settings, these models consistently outperform general-purpose baselines. Notably, they achieve up to a 25% improvement with Best-of-N sampling, while also improving robustness to input noise, enabling effective data filtering, and supporting RL-training of policy models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_11963 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ToolRM: Outcome Reward Models for Tool-Calling Large Language Models Agarwal, Mayank Abdelaziz, Ibrahim Basu, Kinjal Unuvar, Merve Lastras, Luis A. Rizk, Yara Kapanipathi, Pavan Computation and Language As large language models (LLMs) increasingly interact with external tools, reward modeling for tool use has emerged as a critical yet underexplored area of research. Existing reward models, trained primarily on natural language outputs, struggle to evaluate tool-based reasoning and execution. To quantify this gap, we introduce FC-RewardBench, the first benchmark to systematically evaluate reward models in tool-calling scenarios. Our analysis shows that current reward models frequently miss key signals of effective tool use, highlighting the need for domain-specific modeling. We address this by proposing a training framework for outcome reward models using data synthesized from permissively licensed, open-weight LLMs. We introduce ToolRM - a suite of reward models for tool-use ranging from 1.7B to 14B parameters. Across diverse settings, these models consistently outperform general-purpose baselines. Notably, they achieve up to a 25% improvement with Best-of-N sampling, while also improving robustness to input noise, enabling effective data filtering, and supporting RL-training of policy models. |
| title | ToolRM: Outcome Reward Models for Tool-Calling Large Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2509.11963 |