CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866913703427833856 |
|---|---|
| author | Zhang, Alexander Dong, Marcus Liu, Jiaheng Zhang, Wei Wang, Yejie Yang, Jian Zhang, Ge Liu, Tianyu Peng, Zhongyuan Tan, Yingshui Zhang, Yuanxing Wang, Zhexu Wang, Weixun He, Yancheng Deng, Ken Zhou, Wangchunshu Huang, Wenhao Zhang, Zhaoxiang |
| author_facet | Zhang, Alexander Dong, Marcus Liu, Jiaheng Zhang, Wei Wang, Yejie Yang, Jian Zhang, Ge Liu, Tianyu Peng, Zhongyuan Tan, Yingshui Zhang, Yuanxing Wang, Zhexu Wang, Weixun He, Yancheng Deng, Ken Zhou, Wangchunshu Huang, Wenhao Zhang, Zhaoxiang |
| contents | The critique capacity of Large Language Models (LLMs) is essential for reasoning abilities, which can provide necessary suggestions (e.g., detailed analysis and constructive feedback). Therefore, how to evaluate the critique capacity of LLMs has drawn great attention and several critique benchmarks have been proposed. However, existing critique benchmarks usually have the following limitations: (1). Focusing on diverse reasoning tasks in general domains and insufficient evaluation on code tasks (e.g., only covering code generation task), where the difficulty of queries is relatively easy (e.g., the code queries of CriticBench are from Humaneval and MBPP). (2). Lacking comprehensive evaluation from different dimensions. To address these limitations, we introduce a holistic code critique benchmark for LLMs called CodeCriticBench. Specifically, our CodeCriticBench includes two mainstream code tasks (i.e., code generation and code QA) with different difficulties. Besides, the evaluation protocols include basic critique evaluation and advanced critique evaluation for different characteristics, where fine-grained evaluation checklists are well-designed for advanced settings. Finally, we conduct extensive experimental results of existing LLMs, which show the effectiveness of CodeCriticBench. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_16614 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models Zhang, Alexander Dong, Marcus Liu, Jiaheng Zhang, Wei Wang, Yejie Yang, Jian Zhang, Ge Liu, Tianyu Peng, Zhongyuan Tan, Yingshui Zhang, Yuanxing Wang, Zhexu Wang, Weixun He, Yancheng Deng, Ken Zhou, Wangchunshu Huang, Wenhao Zhang, Zhaoxiang Computation and Language The critique capacity of Large Language Models (LLMs) is essential for reasoning abilities, which can provide necessary suggestions (e.g., detailed analysis and constructive feedback). Therefore, how to evaluate the critique capacity of LLMs has drawn great attention and several critique benchmarks have been proposed. However, existing critique benchmarks usually have the following limitations: (1). Focusing on diverse reasoning tasks in general domains and insufficient evaluation on code tasks (e.g., only covering code generation task), where the difficulty of queries is relatively easy (e.g., the code queries of CriticBench are from Humaneval and MBPP). (2). Lacking comprehensive evaluation from different dimensions. To address these limitations, we introduce a holistic code critique benchmark for LLMs called CodeCriticBench. Specifically, our CodeCriticBench includes two mainstream code tasks (i.e., code generation and code QA) with different difficulties. Besides, the evaluation protocols include basic critique evaluation and advanced critique evaluation for different characteristics, where fine-grained evaluation checklists are well-designed for advanced settings. Finally, we conduct extensive experimental results of existing LLMs, which show the effectiveness of CodeCriticBench. |
| title | CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2502.16614 |