Automated Benchmark Generation for Repository-Level Coding Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vergopoulos, Konstantinos, Müller, Mark Niklas, Vechev, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
von: Mündler, Niels, et al.
Veröffentlicht: (2024)
von: Mündler, Niels, et al.
Veröffentlicht: (2024)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
von: Thillen, Alex, et al.
Veröffentlicht: (2026)
On the Impacts of Contexts on Repository-Level Code Generation
von: Hai, Nam Le, et al.
Veröffentlicht: (2024)
von: Hai, Nam Le, et al.
Veröffentlicht: (2024)
A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
von: Lian, Keke, et al.
Veröffentlicht: (2025)
von: Lian, Keke, et al.
Veröffentlicht: (2025)
Beyond Code Snippets: Benchmarking LLMs on Repository-Level Question Answering
von: Alebachew, Yoseph Berhanu, et al.
Veröffentlicht: (2026)
von: Alebachew, Yoseph Berhanu, et al.
Veröffentlicht: (2026)
RepoRepair: Leveraging Code Documentation for Repository-Level Automated Program Repair
von: Pan, Zhongqiang, et al.
Veröffentlicht: (2026)
von: Pan, Zhongqiang, et al.
Veröffentlicht: (2026)
CATCODER: Repository-Level Code Generation with Relevant Code and Type Context
von: Pan, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Pan, Zhiyuan, et al.
Veröffentlicht: (2024)
GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository Leveraging
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
von: Ni, Ziyi, et al.
Veröffentlicht: (2025)
From Laboratory to Real-World Applications: Benchmarking Agentic Code Reasoning at the Repository Level
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
Persistent Cross-Attempt State Optimization for Repository-Level Code Generation
von: Pan, Ruwei, et al.
Veröffentlicht: (2026)
von: Pan, Ruwei, et al.
Veröffentlicht: (2026)
In Line with Context: Repository-Level Code Generation via Context Inlining
von: Hu, Chao, et al.
Veröffentlicht: (2026)
von: Hu, Chao, et al.
Veröffentlicht: (2026)
ReCUBE: Evaluating Repository-Level Context Utilization in Code Generation
von: Hong, Jiseung, et al.
Veröffentlicht: (2026)
von: Hong, Jiseung, et al.
Veröffentlicht: (2026)
Toward Executable Repository-Level Code Generation via Environment Alignment
von: Pan, Ruwei, et al.
Veröffentlicht: (2026)
von: Pan, Ruwei, et al.
Veröffentlicht: (2026)
ToolFuzz -- Automated Agent Tool Testing
von: Milev, Ivan, et al.
Veröffentlicht: (2025)
von: Milev, Ivan, et al.
Veröffentlicht: (2025)
Instruction Tuning for Secure Code Generation
von: He, Jingxuan, et al.
Veröffentlicht: (2024)
von: He, Jingxuan, et al.
Veröffentlicht: (2024)
REPOFUSE: Repository-Level Code Completion with Fused Dual Context
von: Liang, Ming, et al.
Veröffentlicht: (2024)
von: Liang, Ming, et al.
Veröffentlicht: (2024)
ATime-Consistent Benchmark for Repository-Level Software Engineering Evaluation
von: Xianpeng, et al.
Veröffentlicht: (2026)
von: Xianpeng, et al.
Veröffentlicht: (2026)
Toward Scalable Automated Repository-Level Datasets for Software Vulnerability Detection
von: Lbath, Amine
Veröffentlicht: (2026)
von: Lbath, Amine
Veröffentlicht: (2026)
Automated Customization of LLMs for Enterprise Code Repositories Using Semantic Scopes
von: Finkler, Ulrich, et al.
Veröffentlicht: (2026)
von: Finkler, Ulrich, et al.
Veröffentlicht: (2026)
Hierarchical Repository-Level Code Summarization for Business Applications Using Local LLMs
von: Dhulshette, Nilesh, et al.
Veröffentlicht: (2025)
von: Dhulshette, Nilesh, et al.
Veröffentlicht: (2025)
AACR-Bench: Evaluating Automatic Code Review with Holistic Repository-Level Context
von: Zhang, Lei, et al.
Veröffentlicht: (2026)
von: Zhang, Lei, et al.
Veröffentlicht: (2026)
AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code Completion
von: Jiang, Tianyue, et al.
Veröffentlicht: (2026)
von: Jiang, Tianyue, et al.
Veröffentlicht: (2026)
Class-Level Code Generation from Natural Language Using Iterative, Tool-Enhanced Reasoning over Repository
von: Deshpande, Ajinkya, et al.
Veröffentlicht: (2024)
von: Deshpande, Ajinkya, et al.
Veröffentlicht: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
A Study on the Impact of Fault localization Granularity for Repository-Scale Code Repair Tasks
von: Townsend, Joseph, et al.
Veröffentlicht: (2026)
von: Townsend, Joseph, et al.
Veröffentlicht: (2026)
RepoHyper: Search-Expand-Refine on Semantic Graphs for Repository-Level Code Completion
von: Phan, Huy N., et al.
Veröffentlicht: (2024)
von: Phan, Huy N., et al.
Veröffentlicht: (2024)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
von: Wang, Yanli, et al.
Veröffentlicht: (2024)
von: Wang, Yanli, et al.
Veröffentlicht: (2024)
Relative Positioning Based Code Chunking Method For Rich Context Retrieval In Repository Level Code Completion Task With Code Language Model
von: Rahman, Imranur, et al.
Veröffentlicht: (2025)
von: Rahman, Imranur, et al.
Veröffentlicht: (2025)
Beyond Function-Level Search: Repository-Aware Dual-Encoder Code Retrieval with Adversarial Verification
von: Liu, Aofan, et al.
Veröffentlicht: (2025)
von: Liu, Aofan, et al.
Veröffentlicht: (2025)
RepoReviewer: A Local-First Multi-Agent Architecture for Repository-Level Code Review
von: Zhang, Peng
Veröffentlicht: (2026)
von: Zhang, Peng
Veröffentlicht: (2026)
LLM-Assisted Repository-Level Generation with Structured Spec-Driven Engineering
von: Feng, Shuzhao, et al.
Veröffentlicht: (2026)
von: Feng, Shuzhao, et al.
Veröffentlicht: (2026)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
von: Chu, Junjie, et al.
Veröffentlicht: (2026)
von: Chu, Junjie, et al.
Veröffentlicht: (2026)
SWE-QA-Pro: A Representative Benchmark and Scalable Training Recipe for Repository-Level Code Understanding
von: Cai, Songcheng, et al.
Veröffentlicht: (2026)
von: Cai, Songcheng, et al.
Veröffentlicht: (2026)
Needle in the Repo: A Benchmark for Maintainability in AI-Generated Repository Edits
von: Zhu, Haichao, et al.
Veröffentlicht: (2026)
von: Zhu, Haichao, et al.
Veröffentlicht: (2026)
RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models
von: Liu, Jingjing, et al.
Veröffentlicht: (2025)
von: Liu, Jingjing, et al.
Veröffentlicht: (2025)
RTLRepoCoder: Repository-Level RTL Code Completion through the Combination of Fine-Tuning and Retrieval Augmentation
von: Wu, Peiyang, et al.
Veröffentlicht: (2025)
von: Wu, Peiyang, et al.
Veröffentlicht: (2025)
On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing
von: Kovrigin, Alexander, et al.
Veröffentlicht: (2024)
von: Kovrigin, Alexander, et al.
Veröffentlicht: (2024)
Skeleton-Guided-Translation: A Benchmarking Framework for Code Repository Translation with Fine-Grained Quality Evaluation
von: Zhang, Xing, et al.
Veröffentlicht: (2025)
von: Zhang, Xing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026) -
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
von: Mündler, Niels, et al.
Veröffentlicht: (2024) -
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
von: Thillen, Alex, et al.
Veröffentlicht: (2026) -
On the Impacts of Contexts on Repository-Level Code Generation
von: Hai, Nam Le, et al.
Veröffentlicht: (2024) -
A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
von: Lian, Keke, et al.
Veröffentlicht: (2025)