CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xie, Yiqing, Xie, Alex, Sheth, Divyanshu, Liu, Pengfei, Fried, Daniel, Rose, Carolyn
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929525225422848
author Xie, Yiqing
Xie, Alex
Sheth, Divyanshu
Liu, Pengfei
Fried, Daniel
Rose, Carolyn
author_facet Xie, Yiqing
Xie, Alex
Sheth, Divyanshu
Liu, Pengfei
Fried, Daniel
Rose, Carolyn
contents To adequately test modern code generation systems, evaluation benchmarks must execute and test the code generated by the system. However, these execution and testing requirements have largely limited benchmarks to settings where code is easily executable or has human-written tests. To facilitate evaluation of code generation systems across diverse scenarios, we present CodeBenchGen, a framework to create scalable execution-based benchmarks from naturally occurring code sources. Specifically, we leverage a large language model (LLM) to sandbox arbitrary pieces of code into evaluation examples, including test cases for execution-based evaluation. We illustrate the usefulness of our framework by creating a dataset, Exec-CSN, which includes 1,931 examples involving 293 libraries converted from code in 367 GitHub repositories taken from the Code- SearchNet dataset. To demonstrate the solvability of examples in Exec-CSN, we present a human study demonstrating that 81.3% of the examples can be solved by humans and 61% are rated as "requires effort to solve". We conduct code generation experiments on open-source and proprietary models and analyze the performance of both humans and models. We provide code and data at: https://github.com/yiqingxyq/CodeBenchGen.
format Preprint
id arxiv_https___arxiv_org_abs_2404_00566
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
Xie, Yiqing
Xie, Alex
Sheth, Divyanshu
Liu, Pengfei
Fried, Daniel
Rose, Carolyn
Software Engineering
Computation and Language
To adequately test modern code generation systems, evaluation benchmarks must execute and test the code generated by the system. However, these execution and testing requirements have largely limited benchmarks to settings where code is easily executable or has human-written tests. To facilitate evaluation of code generation systems across diverse scenarios, we present CodeBenchGen, a framework to create scalable execution-based benchmarks from naturally occurring code sources. Specifically, we leverage a large language model (LLM) to sandbox arbitrary pieces of code into evaluation examples, including test cases for execution-based evaluation. We illustrate the usefulness of our framework by creating a dataset, Exec-CSN, which includes 1,931 examples involving 293 libraries converted from code in 367 GitHub repositories taken from the Code- SearchNet dataset. To demonstrate the solvability of examples in Exec-CSN, we present a human study demonstrating that 81.3% of the examples can be solved by humans and 61% are rated as "requires effort to solve". We conduct code generation experiments on open-source and proprietary models and analyze the performance of both humans and models. We provide code and data at: https://github.com/yiqingxyq/CodeBenchGen.
title CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
topic Software Engineering
Computation and Language
url https://arxiv.org/abs/2404.00566