CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yuwei, Luo, Ziyang, Tian, Yuchen, Lin, Hongzhan, Yan, Weixiang, Li, Annan, Ma, Jing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916392494694400
author Zhao, Yuwei
Luo, Ziyang
Tian, Yuchen
Lin, Hongzhan
Yan, Weixiang
Li, Annan
Ma, Jing
author_facet Zhao, Yuwei
Luo, Ziyang
Tian, Yuchen
Lin, Hongzhan
Yan, Weixiang
Li, Annan
Ma, Jing
contents Recent advancements in large language models (LLMs) have showcased impressive code generation capabilities, primarily evaluated through language-to-code benchmarks. However, these benchmarks may not fully capture a model's code understanding abilities. We introduce CodeJudge-Eval (CJ-Eval), a novel benchmark designed to assess LLMs' code understanding abilities from the perspective of code judging rather than code generation. CJ-Eval challenges models to determine the correctness of provided code solutions, encompassing various error types and compilation issues. By leveraging a diverse set of problems and a fine-grained judging system, CJ-Eval addresses the limitations of traditional benchmarks, including the potential memorization of solutions. Evaluation of 12 well-known LLMs on CJ-Eval reveals that even state-of-the-art models struggle, highlighting the benchmark's ability to probe deeper into models' code understanding abilities. Our codes and benchmark are available at \url{https://github.com/CodeLLM-Research/CodeJudge-Eval}.
format Preprint
id arxiv_https___arxiv_org_abs_2408_10718
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
Zhao, Yuwei
Luo, Ziyang
Tian, Yuchen
Lin, Hongzhan
Yan, Weixiang
Li, Annan
Ma, Jing
Software Engineering
Computation and Language
Recent advancements in large language models (LLMs) have showcased impressive code generation capabilities, primarily evaluated through language-to-code benchmarks. However, these benchmarks may not fully capture a model's code understanding abilities. We introduce CodeJudge-Eval (CJ-Eval), a novel benchmark designed to assess LLMs' code understanding abilities from the perspective of code judging rather than code generation. CJ-Eval challenges models to determine the correctness of provided code solutions, encompassing various error types and compilation issues. By leveraging a diverse set of problems and a fine-grained judging system, CJ-Eval addresses the limitations of traditional benchmarks, including the potential memorization of solutions. Evaluation of 12 well-known LLMs on CJ-Eval reveals that even state-of-the-art models struggle, highlighting the benchmark's ability to probe deeper into models' code understanding abilities. Our codes and benchmark are available at \url{https://github.com/CodeLLM-Research/CodeJudge-Eval}.
title CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
topic Software Engineering
Computation and Language
url https://arxiv.org/abs/2408.10718