Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Dewu, Wang, Yanlin, Shi, Ensheng, Liu, Xilin, Ma, Yuchi, Zhang, Hongyu, Zheng, Zibin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910878495932416
author Zheng, Dewu
Wang, Yanlin
Shi, Ensheng
Liu, Xilin
Ma, Yuchi
Zhang, Hongyu
Zheng, Zibin
author_facet Zheng, Dewu
Wang, Yanlin
Shi, Ensheng
Liu, Xilin
Ma, Yuchi
Zhang, Hongyu
Zheng, Zibin
contents With the rapid advancement of large language models (LLMs), extensive research has been conducted to investigate the code generation capabilities of LLMs. However, existing efforts primarily focus on general-domain tasks, leaving LLMs' code generation performance in real-world application domains underexplored. This raises a critical question: can a model's general-domain coding ability reliably represent its ability in specialized domains? In this paper, we introduce DomainCodeBench, a multi-domain code generation benchmark designed to systematically evaluate LLMs across 12 software application domains and 15 programming languages. DomainCodeBench contains 2,400 manually verified tasks with ground truth, human-annotated docstrings, and fine-grained dependency information to ensure more coverage of domain-specific challenges. Specifically, we first identify the most popular application domains by topic mining. Then, we curate coding tasks based on commonly used frameworks and platforms in each domain. We obtain several findings through extensive experiments on DomainCodeBench with ten mainstream LLMs. (1) Performance decoupling: experiments reveal that top general-domain models do not consistently excel in specific application domains; (2) Domain-specific weaknesses: LLMs often fail due to domain knowledge gaps and third-party library misusage; (3) Contextual enhancement: we show that augmenting prompts with domain-specific knowledge improves performance by around 38.17%, providing actionable insights for performance optimization. Our replication package, including the benchmark, source code, and experimental results, is available at https://github.com/DeepSoftwareAnalytics/DomainCodeBench.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18573
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
Zheng, Dewu
Wang, Yanlin
Shi, Ensheng
Liu, Xilin
Ma, Yuchi
Zhang, Hongyu
Zheng, Zibin
Software Engineering
Artificial Intelligence
Computation and Language
With the rapid advancement of large language models (LLMs), extensive research has been conducted to investigate the code generation capabilities of LLMs. However, existing efforts primarily focus on general-domain tasks, leaving LLMs' code generation performance in real-world application domains underexplored. This raises a critical question: can a model's general-domain coding ability reliably represent its ability in specialized domains? In this paper, we introduce DomainCodeBench, a multi-domain code generation benchmark designed to systematically evaluate LLMs across 12 software application domains and 15 programming languages. DomainCodeBench contains 2,400 manually verified tasks with ground truth, human-annotated docstrings, and fine-grained dependency information to ensure more coverage of domain-specific challenges. Specifically, we first identify the most popular application domains by topic mining. Then, we curate coding tasks based on commonly used frameworks and platforms in each domain. We obtain several findings through extensive experiments on DomainCodeBench with ten mainstream LLMs. (1) Performance decoupling: experiments reveal that top general-domain models do not consistently excel in specific application domains; (2) Domain-specific weaknesses: LLMs often fail due to domain knowledge gaps and third-party library misusage; (3) Contextual enhancement: we show that augmenting prompts with domain-specific knowledge improves performance by around 38.17%, providing actionable insights for performance optimization. Our replication package, including the benchmark, source code, and experimental results, is available at https://github.com/DeepSoftwareAnalytics/DomainCodeBench.
title Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
topic Software Engineering
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2412.18573