ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lu, Pengrui, Zhang, Shiqi, Hou, Yunzhong, Ye, Lyumanshan, Huang, Chaoyi, Chen, Zixi, Zeng, Ji, Jiang, Hantao, Liu, Pengfei, Wang, Yiwei, Yang, Ming-Hsuan
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914315314921472
author Lu, Pengrui
Zhang, Shiqi
Hou, Yunzhong
Ye, Lyumanshan
Huang, Chaoyi
Chen, Zixi
Zeng, Ji
Jiang, Hantao
Liu, Pengfei
Wang, Yiwei
Yang, Ming-Hsuan
author_facet Lu, Pengrui
Zhang, Shiqi
Hou, Yunzhong
Ye, Lyumanshan
Huang, Chaoyi
Chen, Zixi
Zeng, Ji
Jiang, Hantao
Liu, Pengfei
Wang, Yiwei
Yang, Ming-Hsuan
contents Recent coding agents can generate complete codebases from simple prompts, yet existing evaluations focus on issue-level bug fixing and lag behind end-to-end development. We introduce ProjDevBench, an end-to-end benchmark that provides project requirements to coding agents and evaluates the resulting repositories. Combining Online Judge (OJ) testing with LLM-assisted code review, the benchmark evaluates agents on (1) system architecture design, (2) functional correctness, and (3) iterative solution refinement. We curate 20 programming problems across 8 categories, covering both concept-oriented tasks and real-world application scenarios, and evaluate six coding agents built on different LLM backends. Our evaluation reports an overall acceptance rate of 27.38%: agents handle basic functionality and data structures but struggle with complex system design, time complexity optimization, and resource management. Our benchmark is available at https://github.com/zsworld6/projdevbench.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01655
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
Lu, Pengrui
Zhang, Shiqi
Hou, Yunzhong
Ye, Lyumanshan
Huang, Chaoyi
Chen, Zixi
Zeng, Ji
Jiang, Hantao
Liu, Pengfei
Wang, Yiwei
Yang, Ming-Hsuan
Artificial Intelligence
Software Engineering
Recent coding agents can generate complete codebases from simple prompts, yet existing evaluations focus on issue-level bug fixing and lag behind end-to-end development. We introduce ProjDevBench, an end-to-end benchmark that provides project requirements to coding agents and evaluates the resulting repositories. Combining Online Judge (OJ) testing with LLM-assisted code review, the benchmark evaluates agents on (1) system architecture design, (2) functional correctness, and (3) iterative solution refinement. We curate 20 programming problems across 8 categories, covering both concept-oriented tasks and real-world application scenarios, and evaluate six coding agents built on different LLM backends. Our evaluation reports an overall acceptance rate of 27.38%: agents handle basic functionality and data structures but struggle with complex system design, time complexity optimization, and resource management. Our benchmark is available at https://github.com/zsworld6/projdevbench.
title ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
topic Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2602.01655