CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Liu, Shudong, Liu, Hongwei, Liu, Junnan, Xiao, Linchen, Gao, Songyang, Lyu, Chengqi, Gu, Yuzhe, Zhang, Wenwei, Wong, Derek F., Zhang, Songyang, Chen, Kai
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918115485417472
author Liu, Shudong
Liu, Hongwei
Liu, Junnan
Xiao, Linchen
Gao, Songyang
Lyu, Chengqi
Gu, Yuzhe
Zhang, Wenwei
Wong, Derek F.
Zhang, Songyang
Chen, Kai
author_facet Liu, Shudong
Liu, Hongwei
Liu, Junnan
Xiao, Linchen
Gao, Songyang
Lyu, Chengqi
Gu, Yuzhe
Zhang, Wenwei
Wong, Derek F.
Zhang, Songyang
Chen, Kai
contents Answer verification is crucial not only for evaluating large language models (LLMs) by matching their unstructured outputs against standard answers, but also serves as the reward model to guide LLM optimization. Most evaluation frameworks rely on regularized matching or employ general LLMs for answer verification, which demands extensive, repetitive customization for regex rules or evaluation prompts. Two fundamental limitations persist in current methodologies: 1) the absence of comprehensive benchmarks that systematically evaluate verification capabilities across different LLMs; and 2) the nascent stage of verifier development, where existing approaches lack both the robustness to handle complex edge cases and the generalizability across different domains. In this work, we develop CompassVerifier, an accurate and robust lightweight verifier model for evaluation and outcome reward. It demonstrates multi-domain competency spanning math, knowledge, and diverse reasoning tasks, with the capability to process various answer types, including multi-subproblems, formulas, and sequence answers, while effectively identifying abnormal/invalid responses. We introduce VerifierBench benchmark comprising model outputs collected from multiple data sources, augmented through manual analysis of metaerror patterns to enhance CompassVerifier. We anticipate that CompassVerifier and VerifierBench will facilitate answer verification, evaluation protocols, and reinforcement learning research. Code and dataset are available at https://github.com/open-compass/CompassVerifier.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03686
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
Liu, Shudong
Liu, Hongwei
Liu, Junnan
Xiao, Linchen
Gao, Songyang
Lyu, Chengqi
Gu, Yuzhe
Zhang, Wenwei
Wong, Derek F.
Zhang, Songyang
Chen, Kai
Computation and Language
Artificial Intelligence
Answer verification is crucial not only for evaluating large language models (LLMs) by matching their unstructured outputs against standard answers, but also serves as the reward model to guide LLM optimization. Most evaluation frameworks rely on regularized matching or employ general LLMs for answer verification, which demands extensive, repetitive customization for regex rules or evaluation prompts. Two fundamental limitations persist in current methodologies: 1) the absence of comprehensive benchmarks that systematically evaluate verification capabilities across different LLMs; and 2) the nascent stage of verifier development, where existing approaches lack both the robustness to handle complex edge cases and the generalizability across different domains. In this work, we develop CompassVerifier, an accurate and robust lightweight verifier model for evaluation and outcome reward. It demonstrates multi-domain competency spanning math, knowledge, and diverse reasoning tasks, with the capability to process various answer types, including multi-subproblems, formulas, and sequence answers, while effectively identifying abnormal/invalid responses. We introduce VerifierBench benchmark comprising model outputs collected from multiple data sources, augmented through manual analysis of metaerror patterns to enhance CompassVerifier. We anticipate that CompassVerifier and VerifierBench will facilitate answer verification, evaluation protocols, and reinforcement learning research. Code and dataset are available at https://github.com/open-compass/CompassVerifier.
title CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.03686