CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Sizhe, Wang, Zhengren, Ma, Dongsheng, Yu, Yongan, Ling, Rui, Li, Zhiyu, Xiong, Feiyu, Zhang, Wentao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918446563852288
author Wang, Sizhe
Wang, Zhengren
Ma, Dongsheng
Yu, Yongan
Ling, Rui
Li, Zhiyu
Xiong, Feiyu
Zhang, Wentao
author_facet Wang, Sizhe
Wang, Zhengren
Ma, Dongsheng
Yu, Yongan
Ling, Rui
Li, Zhiyu
Xiong, Feiyu
Zhang, Wentao
contents Modern software development demands code that is maintainable, testable, and scalable by organizing the implementation into modular components with iterative reuse of existing codes. We formalize this iterative, multi-turn paradigm as codeflow and introduce CodeFlowBench, the first benchmark designed to comprehensively evaluate LLMs' ability to perform codeflow - implementing new functionality by reusing existing functions over multiple turns. CodeFlowBench comprises two complementary components: CodeFlowBench-Comp, a core collection of 5,000+ competitive programming problems from Codeforces updated via an automated pipeline and CodeFlowBench-Repo, which is sourced from GitHub repositories to better reflect real-world scenarios. Furthermore, a novel evaluation framework featured dual assessment protocol and structural metrics derived from dependency trees is introduced. Extensive experiments reveal significant performance degradation in multi-turn codeflow scenarios. Furthermore, our in-depth analysis illustrates that model performance inversely correlates with dependency complexity. These findings not only highlight the critical challenges for supporting real-world workflows, but also establish CodeFlowBench as an essential tool for advancing code generation research.
format Preprint
id arxiv_https___arxiv_org_abs_2504_21751
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
Wang, Sizhe
Wang, Zhengren
Ma, Dongsheng
Yu, Yongan
Ling, Rui
Li, Zhiyu
Xiong, Feiyu
Zhang, Wentao
Software Engineering
Computation and Language
Modern software development demands code that is maintainable, testable, and scalable by organizing the implementation into modular components with iterative reuse of existing codes. We formalize this iterative, multi-turn paradigm as codeflow and introduce CodeFlowBench, the first benchmark designed to comprehensively evaluate LLMs' ability to perform codeflow - implementing new functionality by reusing existing functions over multiple turns. CodeFlowBench comprises two complementary components: CodeFlowBench-Comp, a core collection of 5,000+ competitive programming problems from Codeforces updated via an automated pipeline and CodeFlowBench-Repo, which is sourced from GitHub repositories to better reflect real-world scenarios. Furthermore, a novel evaluation framework featured dual assessment protocol and structural metrics derived from dependency trees is introduced. Extensive experiments reveal significant performance degradation in multi-turn codeflow scenarios. Furthermore, our in-depth analysis illustrates that model performance inversely correlates with dependency complexity. These findings not only highlight the critical challenges for supporting real-world workflows, but also establish CodeFlowBench as an essential tool for advancing code generation research.
title CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
topic Software Engineering
Computation and Language
url https://arxiv.org/abs/2504.21751