RepoGenesis: Benchmarking End-to-End Microservice Generation from Readme to Repository

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Zhiyuan, Yin, Xin, Zhao, Pu, Yang, Fangkai, Wang, Lu, Jia, Ran, Chen, Xu, Lin, Qingwei, Rajmohan, Saravan, Zhang, Dongmei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910129812668416
author Peng, Zhiyuan
Yin, Xin
Zhao, Pu
Yang, Fangkai
Wang, Lu
Jia, Ran
Chen, Xu
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
author_facet Peng, Zhiyuan
Yin, Xin
Zhao, Pu
Yang, Fangkai
Wang, Lu
Jia, Ran
Chen, Xu
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
contents Large language models and agents have achieved remarkable progress in code generation. However, existing benchmarks focus on isolated function/class-level generation (e.g., ClassEval) or modifications to existing codebases (e.g., SWE-Bench), neglecting complete microservice repository generation that reflects real-world 0-to-1 development workflows. To bridge this gap, we introduce RepoGenesis, the first multilingual benchmark for repository-level end-to-end web microservice generation, comprising 106 repositories (60 Python, 46 Java) across 18 domains and 11 frameworks, with 1,258 API endpoints and 2,335 test cases verified through a "review-rebuttal" quality assurance process. We evaluate open-source agents (e.g., DeepCode) and commercial IDEs (e.g., Cursor) using Pass@1, API Coverage (AC), and Deployment Success Rate (DSR). Results reveal that despite high AC (up to 73.91%) and DSR (up to 100%), the best-performing system achieves only 23.67% Pass@1 on Python and 21.45% on Java, exposing deficiencies in architectural coherence, dependency management, and cross-file consistency. Notably, GenesisAgent-8B, fine-tuned on RepoGenesis (train), achieves performance comparable to GPT-5 mini, demonstrating the quality of RepoGenesis for advancing microservice generation. We release our benchmark at https://github.com/pzy2000/RepoGenesis.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13943
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RepoGenesis: Benchmarking End-to-End Microservice Generation from Readme to Repository
Peng, Zhiyuan
Yin, Xin
Zhao, Pu
Yang, Fangkai
Wang, Lu
Jia, Ran
Chen, Xu
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
Software Engineering
Large language models and agents have achieved remarkable progress in code generation. However, existing benchmarks focus on isolated function/class-level generation (e.g., ClassEval) or modifications to existing codebases (e.g., SWE-Bench), neglecting complete microservice repository generation that reflects real-world 0-to-1 development workflows. To bridge this gap, we introduce RepoGenesis, the first multilingual benchmark for repository-level end-to-end web microservice generation, comprising 106 repositories (60 Python, 46 Java) across 18 domains and 11 frameworks, with 1,258 API endpoints and 2,335 test cases verified through a "review-rebuttal" quality assurance process. We evaluate open-source agents (e.g., DeepCode) and commercial IDEs (e.g., Cursor) using Pass@1, API Coverage (AC), and Deployment Success Rate (DSR). Results reveal that despite high AC (up to 73.91%) and DSR (up to 100%), the best-performing system achieves only 23.67% Pass@1 on Python and 21.45% on Java, exposing deficiencies in architectural coherence, dependency management, and cross-file consistency. Notably, GenesisAgent-8B, fine-tuned on RepoGenesis (train), achieves performance comparable to GPT-5 mini, demonstrating the quality of RepoGenesis for advancing microservice generation. We release our benchmark at https://github.com/pzy2000/RepoGenesis.
title RepoGenesis: Benchmarking End-to-End Microservice Generation from Readme to Repository
topic Software Engineering
url https://arxiv.org/abs/2601.13943