CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Alexander, Dong, Marcus, Liu, Jiaheng, Zhang, Wei, Wang, Yejie, Yang, Jian, Zhang, Ge, Liu, Tianyu, Peng, Zhongyuan, Tan, Yingshui, Zhang, Yuanxing, Wang, Zhexu, Wang, Weixun, He, Yancheng, Deng, Ken, Zhou, Wangchunshu, Huang, Wenhao, Zhang, Zhaoxiang
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913703427833856
author Zhang, Alexander
Dong, Marcus
Liu, Jiaheng
Zhang, Wei
Wang, Yejie
Yang, Jian
Zhang, Ge
Liu, Tianyu
Peng, Zhongyuan
Tan, Yingshui
Zhang, Yuanxing
Wang, Zhexu
Wang, Weixun
He, Yancheng
Deng, Ken
Zhou, Wangchunshu
Huang, Wenhao
Zhang, Zhaoxiang
author_facet Zhang, Alexander
Dong, Marcus
Liu, Jiaheng
Zhang, Wei
Wang, Yejie
Yang, Jian
Zhang, Ge
Liu, Tianyu
Peng, Zhongyuan
Tan, Yingshui
Zhang, Yuanxing
Wang, Zhexu
Wang, Weixun
He, Yancheng
Deng, Ken
Zhou, Wangchunshu
Huang, Wenhao
Zhang, Zhaoxiang
contents The critique capacity of Large Language Models (LLMs) is essential for reasoning abilities, which can provide necessary suggestions (e.g., detailed analysis and constructive feedback). Therefore, how to evaluate the critique capacity of LLMs has drawn great attention and several critique benchmarks have been proposed. However, existing critique benchmarks usually have the following limitations: (1). Focusing on diverse reasoning tasks in general domains and insufficient evaluation on code tasks (e.g., only covering code generation task), where the difficulty of queries is relatively easy (e.g., the code queries of CriticBench are from Humaneval and MBPP). (2). Lacking comprehensive evaluation from different dimensions. To address these limitations, we introduce a holistic code critique benchmark for LLMs called CodeCriticBench. Specifically, our CodeCriticBench includes two mainstream code tasks (i.e., code generation and code QA) with different difficulties. Besides, the evaluation protocols include basic critique evaluation and advanced critique evaluation for different characteristics, where fine-grained evaluation checklists are well-designed for advanced settings. Finally, we conduct extensive experimental results of existing LLMs, which show the effectiveness of CodeCriticBench.
format Preprint
id arxiv_https___arxiv_org_abs_2502_16614
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models
Zhang, Alexander
Dong, Marcus
Liu, Jiaheng
Zhang, Wei
Wang, Yejie
Yang, Jian
Zhang, Ge
Liu, Tianyu
Peng, Zhongyuan
Tan, Yingshui
Zhang, Yuanxing
Wang, Zhexu
Wang, Weixun
He, Yancheng
Deng, Ken
Zhou, Wangchunshu
Huang, Wenhao
Zhang, Zhaoxiang
Computation and Language
The critique capacity of Large Language Models (LLMs) is essential for reasoning abilities, which can provide necessary suggestions (e.g., detailed analysis and constructive feedback). Therefore, how to evaluate the critique capacity of LLMs has drawn great attention and several critique benchmarks have been proposed. However, existing critique benchmarks usually have the following limitations: (1). Focusing on diverse reasoning tasks in general domains and insufficient evaluation on code tasks (e.g., only covering code generation task), where the difficulty of queries is relatively easy (e.g., the code queries of CriticBench are from Humaneval and MBPP). (2). Lacking comprehensive evaluation from different dimensions. To address these limitations, we introduce a holistic code critique benchmark for LLMs called CodeCriticBench. Specifically, our CodeCriticBench includes two mainstream code tasks (i.e., code generation and code QA) with different difficulties. Besides, the evaluation protocols include basic critique evaluation and advanced critique evaluation for different characteristics, where fine-grained evaluation checklists are well-designed for advanced settings. Finally, we conduct extensive experimental results of existing LLMs, which show the effectiveness of CodeCriticBench.
title CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2502.16614