Multi-Docker-Eval: A `Shovel of the Gold Rush' Benchmark on Automatic Environment Building for Software Engineering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Kelin, Liu, Tianyu, Shang, Zeyu, Ma, Yingwei, Yang, Jian, Liu, Jiaheng, Bian, Kaigui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909957455085568
author Fu, Kelin
Liu, Tianyu
Shang, Zeyu
Ma, Yingwei
Yang, Jian
Liu, Jiaheng
Bian, Kaigui
author_facet Fu, Kelin
Liu, Tianyu
Shang, Zeyu
Ma, Yingwei
Yang, Jian
Liu, Jiaheng
Bian, Kaigui
contents Automated environment configuration is a critical bottleneck in scaling software engineering (SWE) automation. To provide a reliable evaluation standard for this task, we present Multi-Docker-Eval benchmark. It includes 40 real-world repositories spanning 9 programming languages and measures both success in achieving executable states and efficiency under realistic constraints. Our extensive evaluation of state-of-the-art LLMs and agent frameworks reveals key insights: (1) the overall success rate of current models is low (F2P at most 37.7%), with environment construction being the primary bottleneck; (2) model size and reasoning length are not decisive factors, and open-source models like DeepSeek-V3.1 and Kimi-K2 are competitive in both efficiency and effectiveness; (3) agent framework and programming language also have significantly influence on success rate. These findings provide actionable guidelines for building scalable, fully automated SWE pipelines.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06915
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Docker-Eval: A `Shovel of the Gold Rush' Benchmark on Automatic Environment Building for Software Engineering
Fu, Kelin
Liu, Tianyu
Shang, Zeyu
Ma, Yingwei
Yang, Jian
Liu, Jiaheng
Bian, Kaigui
Software Engineering
Automated environment configuration is a critical bottleneck in scaling software engineering (SWE) automation. To provide a reliable evaluation standard for this task, we present Multi-Docker-Eval benchmark. It includes 40 real-world repositories spanning 9 programming languages and measures both success in achieving executable states and efficiency under realistic constraints. Our extensive evaluation of state-of-the-art LLMs and agent frameworks reveals key insights: (1) the overall success rate of current models is low (F2P at most 37.7%), with environment construction being the primary bottleneck; (2) model size and reasoning length are not decisive factors, and open-source models like DeepSeek-V3.1 and Kimi-K2 are competitive in both efficiency and effectiveness; (3) agent framework and programming language also have significantly influence on success rate. These findings provide actionable guidelines for building scalable, fully automated SWE pipelines.
title Multi-Docker-Eval: A `Shovel of the Gold Rush' Benchmark on Automatic Environment Building for Software Engineering
topic Software Engineering
url https://arxiv.org/abs/2512.06915