FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Wei, Zhang, Xin, Guo, Zhongxin, Mao, Shaoguang, Luo, Wen, Peng, Guangyue, Huang, Yangyu, Wang, Houfeng, Li, Scarlett |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
by: Liu, Steven, et al.
Published: (2026)
by: Liu, Steven, et al.
Published: (2026)
MRG-Bench: Evaluating and Exploring the Requirements of Context for Repository-Level Code Generation
by: Li, Haiyang
Published: (2025)
by: Li, Haiyang
Published: (2025)
A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
by: Liu, Shuhan, et al.
Published: (2026)
by: Liu, Shuhan, et al.
Published: (2026)
RepoMod-Bench: A Benchmark for Code Repository Modernization via Implementation-Agnostic Testing
by: Li, Xuefeng, et al.
Published: (2026)
by: Li, Xuefeng, et al.
Published: (2026)
MigrationBench: Repository-Level Code Migration Benchmark from Java 8
by: Liu, Linbo, et al.
Published: (2025)
by: Liu, Linbo, et al.
Published: (2025)
Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation
by: Wang, Chong, et al.
Published: (2024)
by: Wang, Chong, et al.
Published: (2024)
AACR-Bench: Evaluating Automatic Code Review with Holistic Repository-Level Context
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
SolContractEval: A Benchmark for Evaluating Contract-Level Solidity Code Generation
by: Ye, Zhifan, et al.
Published: (2025)
by: Ye, Zhifan, et al.
Published: (2025)
A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
by: Lian, Keke, et al.
Published: (2025)
by: Lian, Keke, et al.
Published: (2025)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
by: Duston, Titouan, et al.
Published: (2025)
by: Duston, Titouan, et al.
Published: (2025)
SolEval: Benchmarking Large Language Models for Repository-level Solidity Code Generation
by: Peng, Zhiyuan, et al.
Published: (2025)
by: Peng, Zhiyuan, et al.
Published: (2025)
NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition
by: Deng, Le, et al.
Published: (2025)
by: Deng, Le, et al.
Published: (2025)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
by: Wang, Yanli, et al.
Published: (2024)
by: Wang, Yanli, et al.
Published: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
Automated Benchmark Generation for Repository-Level Coding Tasks
by: Vergopoulos, Konstantinos, et al.
Published: (2025)
by: Vergopoulos, Konstantinos, et al.
Published: (2025)
Closing the Loop: Universal Repository Representation with RPG-Encoder
by: Luo, Jane, et al.
Published: (2026)
by: Luo, Jane, et al.
Published: (2026)
RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
by: Luo, Jane, et al.
Published: (2025)
by: Luo, Jane, et al.
Published: (2025)
CodeMEM: AST-Guided Adaptive Memory for Repository-Level Iterative Code Generation
by: Wang, Peiding, et al.
Published: (2026)
by: Wang, Peiding, et al.
Published: (2026)
CATCODER: Repository-Level Code Generation with Relevant Code and Type Context
by: Pan, Zhiyuan, et al.
Published: (2024)
by: Pan, Zhiyuan, et al.
Published: (2024)
RepoSummary: Feature-Oriented Summarization and Documentation Generation for Code Repositories
by: Zhu, Yifeng, et al.
Published: (2025)
by: Zhu, Yifeng, et al.
Published: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
RustRepoTrans: Repository-level Code Translation Benchmark Targeting Rust
by: Ou, Guangsheng, et al.
Published: (2024)
by: Ou, Guangsheng, et al.
Published: (2024)
Improving Retrieval-Augmented Code Comment Generation by Retrieving for Generation
by: Lu, Hanzhen, et al.
Published: (2024)
by: Lu, Hanzhen, et al.
Published: (2024)
RLCoder: Reinforcement Learning for Repository-Level Code Completion
by: Wang, Yanlin, et al.
Published: (2024)
by: Wang, Yanlin, et al.
Published: (2024)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
by: Ni, Ziyi, et al.
Published: (2025)
by: Ni, Ziyi, et al.
Published: (2025)
ReCUBE: Evaluating Repository-Level Context Utilization in Code Generation
by: Hong, Jiseung, et al.
Published: (2026)
by: Hong, Jiseung, et al.
Published: (2026)
HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation
by: Zheng, Dewu, et al.
Published: (2024)
by: Zheng, Dewu, et al.
Published: (2024)
From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
FGIT: Fault-Guided Fine-Tuning for Code Generation
by: Fan, Lishui, et al.
Published: (2025)
by: Fan, Lishui, et al.
Published: (2025)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
On the Impacts of Contexts on Repository-Level Code Generation
by: Hai, Nam Le, et al.
Published: (2024)
by: Hai, Nam Le, et al.
Published: (2024)
Do Not Treat Code as Natural Language: Implications for Repository-Level Code Generation and Beyond
by: Le-Anh, Minh, et al.
Published: (2026)
by: Le-Anh, Minh, et al.
Published: (2026)
Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition
by: Liu, Mingwei, et al.
Published: (2026)
by: Liu, Mingwei, et al.
Published: (2026)
QuanBench: Benchmarking Quantum Code Generation with Large Language Models
by: Guo, Xiaoyu, et al.
Published: (2025)
by: Guo, Xiaoyu, et al.
Published: (2025)
RepoScope: Leveraging Call Chain-Aware Multi-View Context for Repository-Level Code Generation
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
by: Guo, Hanyang, et al.
Published: (2025)
by: Guo, Hanyang, et al.
Published: (2025)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
by: Xu, Yisen, et al.
Published: (2026)
by: Xu, Yisen, et al.
Published: (2026)
Beyond Code Snippets: Benchmarking LLMs on Repository-Level Question Answering
by: Alebachew, Yoseph Berhanu, et al.
Published: (2026)
by: Alebachew, Yoseph Berhanu, et al.
Published: (2026)
Retrieval-Augmented Code Generation: A Survey with Focus on Repository-Level Approaches
by: Tao, Yicheng, et al.
Published: (2025)
by: Tao, Yicheng, et al.
Published: (2025)
Similar Items
-
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
by: Liu, Steven, et al.
Published: (2026) -
MRG-Bench: Evaluating and Exploring the Requirements of Context for Repository-Level Code Generation
by: Li, Haiyang
Published: (2025) -
A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
by: Liu, Shuhan, et al.
Published: (2026) -
RepoMod-Bench: A Benchmark for Code Repository Modernization via Implementation-Agnostic Testing
by: Li, Xuefeng, et al.
Published: (2026) -
MigrationBench: Repository-Level Code Migration Benchmark from Java 8
by: Liu, Linbo, et al.
Published: (2025)