CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916392494694400 |
|---|---|
| author | Zhao, Yuwei Luo, Ziyang Tian, Yuchen Lin, Hongzhan Yan, Weixiang Li, Annan Ma, Jing |
| author_facet | Zhao, Yuwei Luo, Ziyang Tian, Yuchen Lin, Hongzhan Yan, Weixiang Li, Annan Ma, Jing |
| contents | Recent advancements in large language models (LLMs) have showcased impressive code generation capabilities, primarily evaluated through language-to-code benchmarks. However, these benchmarks may not fully capture a model's code understanding abilities. We introduce CodeJudge-Eval (CJ-Eval), a novel benchmark designed to assess LLMs' code understanding abilities from the perspective of code judging rather than code generation. CJ-Eval challenges models to determine the correctness of provided code solutions, encompassing various error types and compilation issues. By leveraging a diverse set of problems and a fine-grained judging system, CJ-Eval addresses the limitations of traditional benchmarks, including the potential memorization of solutions. Evaluation of 12 well-known LLMs on CJ-Eval reveals that even state-of-the-art models struggle, highlighting the benchmark's ability to probe deeper into models' code understanding abilities. Our codes and benchmark are available at \url{https://github.com/CodeLLM-Research/CodeJudge-Eval}. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_10718 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding? Zhao, Yuwei Luo, Ziyang Tian, Yuchen Lin, Hongzhan Yan, Weixiang Li, Annan Ma, Jing Software Engineering Computation and Language Recent advancements in large language models (LLMs) have showcased impressive code generation capabilities, primarily evaluated through language-to-code benchmarks. However, these benchmarks may not fully capture a model's code understanding abilities. We introduce CodeJudge-Eval (CJ-Eval), a novel benchmark designed to assess LLMs' code understanding abilities from the perspective of code judging rather than code generation. CJ-Eval challenges models to determine the correctness of provided code solutions, encompassing various error types and compilation issues. By leveraging a diverse set of problems and a fine-grained judging system, CJ-Eval addresses the limitations of traditional benchmarks, including the potential memorization of solutions. Evaluation of 12 well-known LLMs on CJ-Eval reveals that even state-of-the-art models struggle, highlighting the benchmark's ability to probe deeper into models' code understanding abilities. Our codes and benchmark are available at \url{https://github.com/CodeLLM-Research/CodeJudge-Eval}. |
| title | CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding? |
| topic | Software Engineering Computation and Language |
| url | https://arxiv.org/abs/2408.10718 |