mHumanEval -- A Multilingual Benchmark to Evaluate Large Language Models for Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Raihan, Nishat, Anastasopoulos, Antonios, Zampieri, Marcos
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916738975662080
author Raihan, Nishat
Anastasopoulos, Antonios
Zampieri, Marcos
author_facet Raihan, Nishat
Anastasopoulos, Antonios
Zampieri, Marcos
contents Recent advancements in large language models (LLMs) have significantly enhanced code generation from natural language prompts. The HumanEval Benchmark, developed by OpenAI, remains the most widely used code generation benchmark. However, this and other Code LLM benchmarks face critical limitations, particularly in task diversity, test coverage, and linguistic scope. Current evaluations primarily focus on English-to-Python conversion tasks with limited test cases, potentially overestimating model performance. While recent works have addressed test coverage and programming language (PL) diversity, code generation from low-resource language prompts remains largely unexplored. To address this gap, we introduce mHumanEval, an extended benchmark supporting prompts in over 200 natural languages. We employ established machine translation methods to compile the benchmark, coupled with a quality assurance process. Furthermore, we provide expert human translations for 15 diverse natural languages (NLs). We conclude by analyzing the multilingual code generation capabilities of state-of-the-art (SOTA) Code LLMs, offering insights into the current landscape of cross-lingual code generation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15037
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle mHumanEval -- A Multilingual Benchmark to Evaluate Large Language Models for Code Generation
Raihan, Nishat
Anastasopoulos, Antonios
Zampieri, Marcos
Computation and Language
Recent advancements in large language models (LLMs) have significantly enhanced code generation from natural language prompts. The HumanEval Benchmark, developed by OpenAI, remains the most widely used code generation benchmark. However, this and other Code LLM benchmarks face critical limitations, particularly in task diversity, test coverage, and linguistic scope. Current evaluations primarily focus on English-to-Python conversion tasks with limited test cases, potentially overestimating model performance. While recent works have addressed test coverage and programming language (PL) diversity, code generation from low-resource language prompts remains largely unexplored. To address this gap, we introduce mHumanEval, an extended benchmark supporting prompts in over 200 natural languages. We employ established machine translation methods to compile the benchmark, coupled with a quality assurance process. Furthermore, we provide expert human translations for 15 diverse natural languages (NLs). We conclude by analyzing the multilingual code generation capabilities of state-of-the-art (SOTA) Code LLMs, offering insights into the current landscape of cross-lingual code generation.
title mHumanEval -- A Multilingual Benchmark to Evaluate Large Language Models for Code Generation
topic Computation and Language
url https://arxiv.org/abs/2410.15037