A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Shuhan, Zhao, Zhiyi, Hu, Xing, Liu, Kui, Yang, Xiaohu, Xia, Xin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CATCODER: Repository-Level Code Generation with Relevant Code and Type Context
von: Pan, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Pan, Zhiyuan, et al.
Veröffentlicht: (2024)
An Empirical Study of Vulnerable Package Dependencies in LLM Repositories
von: Liu, Shuhan, et al.
Veröffentlicht: (2025)
von: Liu, Shuhan, et al.
Veröffentlicht: (2025)
CREME: Robustness Enhancement of Code LLMs via Layer-Aware Model Editing
von: Liu, Shuhan, et al.
Veröffentlicht: (2025)
von: Liu, Shuhan, et al.
Veröffentlicht: (2025)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
von: Pan, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Pan, Zhiyuan, et al.
Veröffentlicht: (2025)
Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition
von: Liu, Mingwei, et al.
Veröffentlicht: (2026)
von: Liu, Mingwei, et al.
Veröffentlicht: (2026)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
Automated Benchmark Generation for Repository-Level Coding Tasks
von: Vergopoulos, Konstantinos, et al.
Veröffentlicht: (2025)
von: Vergopoulos, Konstantinos, et al.
Veröffentlicht: (2025)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
SolContractEval: A Benchmark for Evaluating Contract-Level Solidity Code Generation
von: Ye, Zhifan, et al.
Veröffentlicht: (2025)
von: Ye, Zhifan, et al.
Veröffentlicht: (2025)
MigrationBench: Repository-Level Code Migration Benchmark from Java 8
von: Liu, Linbo, et al.
Veröffentlicht: (2025)
von: Liu, Linbo, et al.
Veröffentlicht: (2025)
Dependency-Guided Repository-Level C-to-Rust Translation with Reinforcement Alignment
von: Feng, Jia, et al.
Veröffentlicht: (2026)
von: Feng, Jia, et al.
Veröffentlicht: (2026)
NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition
von: Deng, Le, et al.
Veröffentlicht: (2025)
von: Deng, Le, et al.
Veröffentlicht: (2025)
Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation
von: Wang, Chong, et al.
Veröffentlicht: (2024)
von: Wang, Chong, et al.
Veröffentlicht: (2024)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
Understanding Practitioners' Expectations on Clear Code Review Comments
von: Chen, Junkai, et al.
Veröffentlicht: (2024)
von: Chen, Junkai, et al.
Veröffentlicht: (2024)
A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
von: Lian, Keke, et al.
Veröffentlicht: (2025)
von: Lian, Keke, et al.
Veröffentlicht: (2025)
Every Maintenance Has Its Exemplar: The Future of Software Maintenance through Migration
von: Chen, Zirui, et al.
Veröffentlicht: (2026)
von: Chen, Zirui, et al.
Veröffentlicht: (2026)
Automated Unit Test Refactoring
von: Gao, Yi, et al.
Veröffentlicht: (2024)
von: Gao, Yi, et al.
Veröffentlicht: (2024)
An Empirical Study of Retrieval-Augmented Code Generation: Challenges and Opportunities
von: Yang, Zezhou, et al.
Veröffentlicht: (2025)
von: Yang, Zezhou, et al.
Veröffentlicht: (2025)
Pre-training by Predicting Program Dependencies for Vulnerability Analysis Tasks
von: Liu, Zhongxin, et al.
Veröffentlicht: (2024)
von: Liu, Zhongxin, et al.
Veröffentlicht: (2024)
RustRepoTrans: Repository-level Code Translation Benchmark Targeting Rust
von: Ou, Guangsheng, et al.
Veröffentlicht: (2024)
von: Ou, Guangsheng, et al.
Veröffentlicht: (2024)
From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
von: Xu, Yisen, et al.
Veröffentlicht: (2026)
von: Xu, Yisen, et al.
Veröffentlicht: (2026)
Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks
von: Hu, Xing, et al.
Veröffentlicht: (2025)
von: Hu, Xing, et al.
Veröffentlicht: (2025)
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
Fight Fire with Fire: How Much Can We Trust ChatGPT on Source Code-Related Tasks?
von: Yu, Xiao, et al.
Veröffentlicht: (2024)
von: Yu, Xiao, et al.
Veröffentlicht: (2024)
CodeMEM: AST-Guided Adaptive Memory for Repository-Level Iterative Code Generation
von: Wang, Peiding, et al.
Veröffentlicht: (2026)
von: Wang, Peiding, et al.
Veröffentlicht: (2026)
Similar but Patched Code Considered Harmful -- The Impact of Similar but Patched Code on Recurring Vulnerability Detection and How to Remove Them
von: Tan, Zixuan, et al.
Veröffentlicht: (2024)
von: Tan, Zixuan, et al.
Veröffentlicht: (2024)
Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
von: Wang, Shufan, et al.
Veröffentlicht: (2025)
von: Wang, Shufan, et al.
Veröffentlicht: (2025)
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
SolEval: Benchmarking Large Language Models for Repository-level Solidity Code Generation
von: Peng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Peng, Zhiyuan, et al.
Veröffentlicht: (2025)
Generating Mitigations for Downstream Projects to Neutralize Upstream Library Vulnerability
von: Chen, Zirui, et al.
Veröffentlicht: (2025)
von: Chen, Zirui, et al.
Veröffentlicht: (2025)
A Rule-Based Approach for UI Migration from Android to iOS
von: Gao, Yi, et al.
Veröffentlicht: (2024)
von: Gao, Yi, et al.
Veröffentlicht: (2024)
DepRadar: Agentic Coordination for Context Aware Defect Impact Analysis in Deep Learning Libraries
von: Gao, Yi, et al.
Veröffentlicht: (2026)
von: Gao, Yi, et al.
Veröffentlicht: (2026)
MRG-Bench: Evaluating and Exploring the Requirements of Context for Repository-Level Code Generation
von: Li, Haiyang
Veröffentlicht: (2025)
von: Li, Haiyang
Veröffentlicht: (2025)
TimeMachine-bench: A Benchmark for Evaluating Model Capabilities in Repository-Level Migration Tasks
von: Fujii, Ryo, et al.
Veröffentlicht: (2026)
von: Fujii, Ryo, et al.
Veröffentlicht: (2026)
Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling
von: Yang, Xu, et al.
Veröffentlicht: (2025)
von: Yang, Xu, et al.
Veröffentlicht: (2025)
RepoScope: Leveraging Call Chain-Aware Multi-View Context for Repository-Level Code Generation
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
RepoTransAgent: Multi-Agent LLM Framework for Repository-Aware Code Translation
von: Guan, Ziqi, et al.
Veröffentlicht: (2025)
von: Guan, Ziqi, et al.
Veröffentlicht: (2025)
Instructive Code Retriever: Learn from Large Language Model's Feedback for Code Intelligence Tasks
von: Lu, Jiawei, et al.
Veröffentlicht: (2024)
von: Lu, Jiawei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CATCODER: Repository-Level Code Generation with Relevant Code and Type Context
von: Pan, Zhiyuan, et al.
Veröffentlicht: (2024) -
An Empirical Study of Vulnerable Package Dependencies in LLM Repositories
von: Liu, Shuhan, et al.
Veröffentlicht: (2025) -
CREME: Robustness Enhancement of Code LLMs via Layer-Aware Model Editing
von: Liu, Shuhan, et al.
Veröffentlicht: (2025) -
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
von: Pan, Zhiyuan, et al.
Veröffentlicht: (2025) -
Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition
von: Liu, Mingwei, et al.
Veröffentlicht: (2026)