Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Jiasheng, Cao, Boxi, Ma, Zhengzhao, Pan, Ruotong, Lin, Hongyu, Lu, Yaojie, Han, Xianpei, Sun, Le
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916427324194816
author Zheng, Jiasheng
Cao, Boxi
Ma, Zhengzhao
Pan, Ruotong
Lin, Hongyu
Lu, Yaojie
Han, Xianpei
Sun, Le
author_facet Zheng, Jiasheng
Cao, Boxi
Ma, Zhengzhao
Pan, Ruotong
Lin, Hongyu
Lu, Yaojie
Han, Xianpei
Sun, Le
contents In recent years, researchers have proposed numerous benchmarks to evaluate the impressive coding capabilities of large language models (LLMs). However, current benchmarks primarily assess the accuracy of LLM-generated code, while neglecting other critical dimensions that also significantly impact code quality in real-world development. Moreover, relying exclusively on correctness as the guiding metric renders LLMs susceptible to data contamination. Therefore, this paper proposes the RACE benchmark, which comprehensively evaluates the quality of code generated by LLMs across 4 dimensions: Readability, mAintainability, Correctness, and Efficiency. Specifically, considering the demand-dependent nature of dimensions beyond correctness, we design various types of user requirements for each dimension to assess the model's ability to generate correct code that also meets user demands. We analyze 28 representative LLMs based on RACE and find that: 1) current correctness-centric benchmarks fail to capture the multifaceted requirements of code in real-world scenarios, while RACE provides a comprehensive evaluation that reveals the defects of LLMs across multiple dimensions; 2) the RACE benchmark serves as an effective tool for resisting the risk of data contamination; 3) even the most advanced code LLMs still encounter significant challenges in customized requirements involving complex instructions; 4) most LLMs exhibit an inherent preference for specific coding style. These findings highlight the need for a multidimensional evaluation of code LLMs, emphasizing metrics beyond correctness for real-world applications. Future efforts should aim to develop novel learning algorithms to enhance code generation under varied constraints and improve coverage and usability for diverse user needs.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11470
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models
Zheng, Jiasheng
Cao, Boxi
Ma, Zhengzhao
Pan, Ruotong
Lin, Hongyu
Lu, Yaojie
Han, Xianpei
Sun, Le
Software Engineering
Artificial Intelligence
Computation and Language
In recent years, researchers have proposed numerous benchmarks to evaluate the impressive coding capabilities of large language models (LLMs). However, current benchmarks primarily assess the accuracy of LLM-generated code, while neglecting other critical dimensions that also significantly impact code quality in real-world development. Moreover, relying exclusively on correctness as the guiding metric renders LLMs susceptible to data contamination. Therefore, this paper proposes the RACE benchmark, which comprehensively evaluates the quality of code generated by LLMs across 4 dimensions: Readability, mAintainability, Correctness, and Efficiency. Specifically, considering the demand-dependent nature of dimensions beyond correctness, we design various types of user requirements for each dimension to assess the model's ability to generate correct code that also meets user demands. We analyze 28 representative LLMs based on RACE and find that: 1) current correctness-centric benchmarks fail to capture the multifaceted requirements of code in real-world scenarios, while RACE provides a comprehensive evaluation that reveals the defects of LLMs across multiple dimensions; 2) the RACE benchmark serves as an effective tool for resisting the risk of data contamination; 3) even the most advanced code LLMs still encounter significant challenges in customized requirements involving complex instructions; 4) most LLMs exhibit an inherent preference for specific coding style. These findings highlight the need for a multidimensional evaluation of code LLMs, emphasizing metrics beyond correctness for real-world applications. Future efforts should aim to develop novel learning algorithms to enhance code generation under varied constraints and improve coverage and usability for diverse user needs.
title Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2407.11470