EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated Code

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qing, Yuhao, Zhu, Boyu, Du, Mingzhe, Guo, Zhijiang, Zhuo, Terry Yue, Zhang, Qianru, Zhang, Jie M., Cui, Heming, Yiu, Siu-Ming, Huang, Dong, Ng, See-Kiong, Tuan, Luu Anh
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916744073838592
author Qing, Yuhao
Zhu, Boyu
Du, Mingzhe
Guo, Zhijiang
Zhuo, Terry Yue
Zhang, Qianru
Zhang, Jie M.
Cui, Heming
Yiu, Siu-Ming
Huang, Dong
Ng, See-Kiong
Tuan, Luu Anh
author_facet Qing, Yuhao
Zhu, Boyu
Du, Mingzhe
Guo, Zhijiang
Zhuo, Terry Yue
Zhang, Qianru
Zhang, Jie M.
Cui, Heming
Yiu, Siu-Ming
Huang, Dong
Ng, See-Kiong
Tuan, Luu Anh
contents Existing code generation benchmarks primarily evaluate functional correctness, with limited focus on code efficiency and often restricted to a single language like Python. To address this gap, we introduce EffiBench-X, the first multi-language benchmark designed to measure the efficiency of LLM-generated code. EffiBench-X supports Python, C++, Java, JavaScript, Ruby, and Golang. It comprises competitive programming tasks with human-expert solutions as efficiency baselines. Evaluating state-of-the-art LLMs on EffiBench-X reveals that while models generate functionally correct code, they consistently underperform human experts in efficiency. Even the most efficient LLM-generated solutions (Qwen3-32B) achieve only around \textbf{62\%} of human efficiency on average, with significant language-specific variations. LLMs show better efficiency in Python, Ruby, and JavaScript than in Java, C++, and Golang. For instance, DeepSeek-R1's Python code is significantly more efficient than its Java code. These results highlight the critical need for research into LLM optimization techniques to improve code efficiency across diverse languages. The dataset and evaluation infrastructure are submitted and available at https://github.com/EffiBench/EffiBench-X.git and https://huggingface.co/datasets/EffiBench/effibench-x.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13004
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated Code
Qing, Yuhao
Zhu, Boyu
Du, Mingzhe
Guo, Zhijiang
Zhuo, Terry Yue
Zhang, Qianru
Zhang, Jie M.
Cui, Heming
Yiu, Siu-Ming
Huang, Dong
Ng, See-Kiong
Tuan, Luu Anh
Computation and Language
Existing code generation benchmarks primarily evaluate functional correctness, with limited focus on code efficiency and often restricted to a single language like Python. To address this gap, we introduce EffiBench-X, the first multi-language benchmark designed to measure the efficiency of LLM-generated code. EffiBench-X supports Python, C++, Java, JavaScript, Ruby, and Golang. It comprises competitive programming tasks with human-expert solutions as efficiency baselines. Evaluating state-of-the-art LLMs on EffiBench-X reveals that while models generate functionally correct code, they consistently underperform human experts in efficiency. Even the most efficient LLM-generated solutions (Qwen3-32B) achieve only around \textbf{62\%} of human efficiency on average, with significant language-specific variations. LLMs show better efficiency in Python, Ruby, and JavaScript than in Java, C++, and Golang. For instance, DeepSeek-R1's Python code is significantly more efficient than its Java code. These results highlight the critical need for research into LLM optimization techniques to improve code efficiency across diverse languages. The dataset and evaluation infrastructure are submitted and available at https://github.com/EffiBench/EffiBench-X.git and https://huggingface.co/datasets/EffiBench/effibench-x.
title EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated Code
topic Computation and Language
url https://arxiv.org/abs/2505.13004