BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zehua, Bajaj, Ati Priya, Handa, Divij, Liu, Siyu, Raj, Arvind S, Chen, Hongkai, Wang, Hulin, Liu, Yibo, Basque, Zion Leonahenahe, Nath, Souradip, Juneja, Vishal, Chapre, Nikhil, Shoshitaishvili, Yan, Doupé, Adam, Baral, Chitta, Wang, Ruoyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915523002892288
author Zhang, Zehua
Bajaj, Ati Priya
Handa, Divij
Liu, Siyu
Raj, Arvind S
Chen, Hongkai
Wang, Hulin
Liu, Yibo
Basque, Zion Leonahenahe
Nath, Souradip
Juneja, Vishal
Chapre, Nikhil
Shoshitaishvili, Yan
Doupé, Adam
Baral, Chitta
Wang, Ruoyu
author_facet Zhang, Zehua
Bajaj, Ati Priya
Handa, Divij
Liu, Siyu
Raj, Arvind S
Chen, Hongkai
Wang, Hulin
Liu, Yibo
Basque, Zion Leonahenahe
Nath, Souradip
Juneja, Vishal
Chapre, Nikhil
Shoshitaishvili, Yan
Doupé, Adam
Baral, Chitta
Wang, Ruoyu
contents Automatically compiling open-source software (OSS) projects is a vital, labor-intensive, and complex task, which makes it a good challenge for LLM Agents. Existing methods rely on manually curated rules and workflows, which cannot adapt to OSS that requires customized configuration or environment setup. Recent attempts using Large Language Models (LLMs) used selective evaluation on a subset of highly rated OSS, a practice that underestimates the realistic challenges of OSS compilation. In practice, compilation instructions are often absent, dependencies are undocumented, and successful builds may even require patching source files or modifying build scripts. We propose a more challenging and realistic benchmark, BUILD-BENCH, comprising OSS that are more diverse in quality, scale, and characteristics. Furthermore, we propose a strong baseline LLM-based agent, OSS-BUILD-AGENT, an effective system with enhanced build instruction retrieval module that achieves state-of-the-art performance on BUILD-BENCH and is adaptable to heterogeneous OSS characteristics. We also provide detailed analysis regarding different compilation method design choices and their influence to the whole task, offering insights to guide future advances. We believe performance on BUILD-BENCH can faithfully reflect an agent's ability to tackle compilation as a complex software engineering tasks, and, as such, our benchmark will spur innovation with a significant impact on downstream applications in the fields of software development and software security.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25248
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
Zhang, Zehua
Bajaj, Ati Priya
Handa, Divij
Liu, Siyu
Raj, Arvind S
Chen, Hongkai
Wang, Hulin
Liu, Yibo
Basque, Zion Leonahenahe
Nath, Souradip
Juneja, Vishal
Chapre, Nikhil
Shoshitaishvili, Yan
Doupé, Adam
Baral, Chitta
Wang, Ruoyu
Software Engineering
Artificial Intelligence
Programming Languages
Automatically compiling open-source software (OSS) projects is a vital, labor-intensive, and complex task, which makes it a good challenge for LLM Agents. Existing methods rely on manually curated rules and workflows, which cannot adapt to OSS that requires customized configuration or environment setup. Recent attempts using Large Language Models (LLMs) used selective evaluation on a subset of highly rated OSS, a practice that underestimates the realistic challenges of OSS compilation. In practice, compilation instructions are often absent, dependencies are undocumented, and successful builds may even require patching source files or modifying build scripts. We propose a more challenging and realistic benchmark, BUILD-BENCH, comprising OSS that are more diverse in quality, scale, and characteristics. Furthermore, we propose a strong baseline LLM-based agent, OSS-BUILD-AGENT, an effective system with enhanced build instruction retrieval module that achieves state-of-the-art performance on BUILD-BENCH and is adaptable to heterogeneous OSS characteristics. We also provide detailed analysis regarding different compilation method design choices and their influence to the whole task, offering insights to guide future advances. We believe performance on BUILD-BENCH can faithfully reflect an agent's ability to tackle compilation as a complex software engineering tasks, and, as such, our benchmark will spur innovation with a significant impact on downstream applications in the fields of software development and software security.
title BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
topic Software Engineering
Artificial Intelligence
Programming Languages
url https://arxiv.org/abs/2509.25248