VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Yuchen, Jiang, Jin, Ren, Zhenbang, Li, Yijun, Cai, Xudong, Liu, Yang, Xu, Xin, Zhang, Mengdi, Shao, Jian, Shen, Yongliang, Xiao, Jun, Zhuang, Yueting
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914335725453312
author Yan, Yuchen
Jiang, Jin
Ren, Zhenbang
Li, Yijun
Cai, Xudong
Liu, Yang
Xu, Xin
Zhang, Mengdi
Shao, Jian
Shen, Yongliang
Xiao, Jun
Zhuang, Yueting
author_facet Yan, Yuchen
Jiang, Jin
Ren, Zhenbang
Li, Yijun
Cai, Xudong
Liu, Yang
Xu, Xin
Zhang, Mengdi
Shao, Jian
Shen, Yongliang
Xiao, Jun
Zhuang, Yueting
contents Large reasoning models such as OpenAI o1 and DeepSeek-R1 have demonstrated remarkable performance in complex reasoning tasks. A critical component of their training is the incorporation of reference-based reward systems within reinforcement learning (RL), where model outputs are evaluated against ground truth references. However, existing reward benchmarks focus on preference comparisons between responses rather than evaluating verification against ground truth references, leaving a critical gap in our ability to evaluate verification systems used in reasoning model training. In this paper, we introduce VerifyBench and its challenging variant VerifyBench-Hard, two benchmarks specifically designed to assess reference-based reward systems. These benchmarks are constructed through meticulous data collection and curation, followed by careful human annotation to ensure high quality. Our comprehensive evaluation reveals that while larger model-based verifiers show promise on standard cases, all current systems demonstrate substantial room for improvement on challenging instances. Through systematic analysis of performance patterns across reasoning tasks and error categories, we provide insights for advancing reference-based reward systems. These benchmarks establish a standardized framework for improving verification accuracy, ultimately enhancing reasoning capabilities in models trained via RL.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15801
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
Yan, Yuchen
Jiang, Jin
Ren, Zhenbang
Li, Yijun
Cai, Xudong
Liu, Yang
Xu, Xin
Zhang, Mengdi
Shao, Jian
Shen, Yongliang
Xiao, Jun
Zhuang, Yueting
Computation and Language
Artificial Intelligence
Large reasoning models such as OpenAI o1 and DeepSeek-R1 have demonstrated remarkable performance in complex reasoning tasks. A critical component of their training is the incorporation of reference-based reward systems within reinforcement learning (RL), where model outputs are evaluated against ground truth references. However, existing reward benchmarks focus on preference comparisons between responses rather than evaluating verification against ground truth references, leaving a critical gap in our ability to evaluate verification systems used in reasoning model training. In this paper, we introduce VerifyBench and its challenging variant VerifyBench-Hard, two benchmarks specifically designed to assess reference-based reward systems. These benchmarks are constructed through meticulous data collection and curation, followed by careful human annotation to ensure high quality. Our comprehensive evaluation reveals that while larger model-based verifiers show promise on standard cases, all current systems demonstrate substantial room for improvement on challenging instances. Through systematic analysis of performance patterns across reasoning tasks and error categories, we provide insights for advancing reference-based reward systems. These benchmarks establish a standardized framework for improving verification accuracy, ultimately enhancing reasoning capabilities in models trained via RL.
title VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.15801