ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866914315314921472 |
|---|---|
| author | Lu, Pengrui Zhang, Shiqi Hou, Yunzhong Ye, Lyumanshan Huang, Chaoyi Chen, Zixi Zeng, Ji Jiang, Hantao Liu, Pengfei Wang, Yiwei Yang, Ming-Hsuan |
| author_facet | Lu, Pengrui Zhang, Shiqi Hou, Yunzhong Ye, Lyumanshan Huang, Chaoyi Chen, Zixi Zeng, Ji Jiang, Hantao Liu, Pengfei Wang, Yiwei Yang, Ming-Hsuan |
| contents | Recent coding agents can generate complete codebases from simple prompts, yet existing evaluations focus on issue-level bug fixing and lag behind end-to-end development. We introduce ProjDevBench, an end-to-end benchmark that provides project requirements to coding agents and evaluates the resulting repositories. Combining Online Judge (OJ) testing with LLM-assisted code review, the benchmark evaluates agents on (1) system architecture design, (2) functional correctness, and (3) iterative solution refinement. We curate 20 programming problems across 8 categories, covering both concept-oriented tasks and real-world application scenarios, and evaluate six coding agents built on different LLM backends. Our evaluation reports an overall acceptance rate of 27.38%: agents handle basic functionality and data structures but struggle with complex system design, time complexity optimization, and resource management. Our benchmark is available at https://github.com/zsworld6/projdevbench. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_01655 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development Lu, Pengrui Zhang, Shiqi Hou, Yunzhong Ye, Lyumanshan Huang, Chaoyi Chen, Zixi Zeng, Ji Jiang, Hantao Liu, Pengfei Wang, Yiwei Yang, Ming-Hsuan Artificial Intelligence Software Engineering Recent coding agents can generate complete codebases from simple prompts, yet existing evaluations focus on issue-level bug fixing and lag behind end-to-end development. We introduce ProjDevBench, an end-to-end benchmark that provides project requirements to coding agents and evaluates the resulting repositories. Combining Online Judge (OJ) testing with LLM-assisted code review, the benchmark evaluates agents on (1) system architecture design, (2) functional correctness, and (3) iterative solution refinement. We curate 20 programming problems across 8 categories, covering both concept-oriented tasks and real-world application scenarios, and evaluate six coding agents built on different LLM backends. Our evaluation reports an overall acceptance rate of 27.38%: agents handle basic functionality and data structures but struggle with complex system design, time complexity optimization, and resource management. Our benchmark is available at https://github.com/zsworld6/projdevbench. |
| title | ProjDevBench: Benchmarking AI Coding Agents on End-to-End Project Development |
| topic | Artificial Intelligence Software Engineering |
| url | https://arxiv.org/abs/2602.01655 |