EpiCoder: Encompassing Diversity and Complexity in Code Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yaoxiang, Li, Haoling, Zhang, Xin, Wu, Jie, Liu, Xiao, Hu, Wenxiang, Guo, Zhongxin, Huang, Yangyu, Xin, Ying, Yang, Yujiu, Su, Jinsong, Chen, Qi, Li, Scarlett
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909831865040896
author Wang, Yaoxiang
Li, Haoling
Zhang, Xin
Wu, Jie
Liu, Xiao
Hu, Wenxiang
Guo, Zhongxin
Huang, Yangyu
Xin, Ying
Yang, Yujiu
Su, Jinsong
Chen, Qi
Li, Scarlett
author_facet Wang, Yaoxiang
Li, Haoling
Zhang, Xin
Wu, Jie
Liu, Xiao
Hu, Wenxiang
Guo, Zhongxin
Huang, Yangyu
Xin, Ying
Yang, Yujiu
Su, Jinsong
Chen, Qi
Li, Scarlett
contents Existing methods for code generation use code snippets as seed data, restricting the complexity and diversity of the synthesized data. In this paper, we introduce a novel feature tree-based synthesis framework, which revolves around hierarchical code features derived from high-level abstractions of code. The feature tree is constructed from raw data and refined iteratively to increase the quantity and diversity of the extracted features, which captures and recognizes more complex patterns and relationships within the code. By adjusting the depth and breadth of the sampled subtrees, our framework provides precise control over the complexity of the generated code, enabling functionalities that range from function-level operations to multi-file scenarios. We fine-tuned widely-used base models to obtain EpiCoder series, achieving state-of-the-art performance on multiple benchmarks at both the function and file levels. In particular, empirical evidence indicates that our approach shows significant potential in the synthesizing of repository-level code data. Our code and data are publicly available at https://github.com/microsoft/EpiCoder.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04694
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EpiCoder: Encompassing Diversity and Complexity in Code Generation
Wang, Yaoxiang
Li, Haoling
Zhang, Xin
Wu, Jie
Liu, Xiao
Hu, Wenxiang
Guo, Zhongxin
Huang, Yangyu
Xin, Ying
Yang, Yujiu
Su, Jinsong
Chen, Qi
Li, Scarlett
Computation and Language
Artificial Intelligence
Existing methods for code generation use code snippets as seed data, restricting the complexity and diversity of the synthesized data. In this paper, we introduce a novel feature tree-based synthesis framework, which revolves around hierarchical code features derived from high-level abstractions of code. The feature tree is constructed from raw data and refined iteratively to increase the quantity and diversity of the extracted features, which captures and recognizes more complex patterns and relationships within the code. By adjusting the depth and breadth of the sampled subtrees, our framework provides precise control over the complexity of the generated code, enabling functionalities that range from function-level operations to multi-file scenarios. We fine-tuned widely-used base models to obtain EpiCoder series, achieving state-of-the-art performance on multiple benchmarks at both the function and file levels. In particular, empirical evidence indicates that our approach shows significant potential in the synthesizing of repository-level code data. Our code and data are publicly available at https://github.com/microsoft/EpiCoder.
title EpiCoder: Encompassing Diversity and Complexity in Code Generation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2501.04694