RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Zhengyang, Li, Ziniu, Xiao, Zhenyang, Ding, Tian, Sun, Ruoyu, Wang, Benyou, Liu, Dayiheng, Huang, Fei, Liu, Tianyu, Yu, Bowen, Lin, Junyang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917901883146240
author Tang, Zhengyang
Li, Ziniu
Xiao, Zhenyang
Ding, Tian
Sun, Ruoyu
Wang, Benyou
Liu, Dayiheng
Huang, Fei
Liu, Tianyu
Yu, Bowen
Lin, Junyang
author_facet Tang, Zhengyang
Li, Ziniu
Xiao, Zhenyang
Ding, Tian
Sun, Ruoyu
Wang, Benyou
Liu, Dayiheng
Huang, Fei
Liu, Tianyu
Yu, Bowen
Lin, Junyang
contents Critiques are important for enhancing the performance of Large Language Models (LLMs), enabling both self-improvement and constructive feedback for others by identifying flaws and suggesting improvements. However, evaluating the critique capabilities of LLMs presents a significant challenge due to the open-ended nature of the task. In this work, we introduce a new benchmark designed to assess the critique capabilities of LLMs. Unlike existing benchmarks, which typically function in an open-loop fashion, our approach employs a closed-loop methodology that evaluates the quality of corrections generated from critiques. Moreover, the benchmark incorporates features such as self-critique, cross-critique, and iterative critique, which are crucial for distinguishing the abilities of advanced reasoning models from more classical ones. We implement this benchmark using eight challenging reasoning tasks. We have several interesting findings. First, despite demonstrating comparable performance in direct chain-of-thought generation, classical LLMs significantly lag behind the advanced reasoning-based model o1-mini across all critique scenarios. Second, in self-critique and iterative critique settings, classical LLMs may even underperform relative to their baseline capabilities. We hope that this benchmark will serve as a valuable resource to guide future advancements. The code and data are available at \url{https://github.com/tangzhy/RealCritic}.
format Preprint
id arxiv_https___arxiv_org_abs_2501_14492
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques
Tang, Zhengyang
Li, Ziniu
Xiao, Zhenyang
Ding, Tian
Sun, Ruoyu
Wang, Benyou
Liu, Dayiheng
Huang, Fei
Liu, Tianyu
Yu, Bowen
Lin, Junyang
Computation and Language
Artificial Intelligence
Machine Learning
Critiques are important for enhancing the performance of Large Language Models (LLMs), enabling both self-improvement and constructive feedback for others by identifying flaws and suggesting improvements. However, evaluating the critique capabilities of LLMs presents a significant challenge due to the open-ended nature of the task. In this work, we introduce a new benchmark designed to assess the critique capabilities of LLMs. Unlike existing benchmarks, which typically function in an open-loop fashion, our approach employs a closed-loop methodology that evaluates the quality of corrections generated from critiques. Moreover, the benchmark incorporates features such as self-critique, cross-critique, and iterative critique, which are crucial for distinguishing the abilities of advanced reasoning models from more classical ones. We implement this benchmark using eight challenging reasoning tasks. We have several interesting findings. First, despite demonstrating comparable performance in direct chain-of-thought generation, classical LLMs significantly lag behind the advanced reasoning-based model o1-mini across all critique scenarios. Second, in self-critique and iterative critique settings, classical LLMs may even underperform relative to their baseline capabilities. We hope that this benchmark will serve as a valuable resource to guide future advancements. The code and data are available at \url{https://github.com/tangzhy/RealCritic}.
title RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2501.14492