EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Wentao, Wang, Jianfeng, Liang, Liheng, Zhao, Yilei, Wen, HaiBin, Zhao, Zhe |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation
par: Liang, Linxi, et autres
Publié: (2025)
par: Liang, Linxi, et autres
Publié: (2025)
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
par: Xia, Chunqiu Steven, et autres
Publié: (2024)
par: Xia, Chunqiu Steven, et autres
Publié: (2024)
HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation
par: Zheng, Dewu, et autres
Publié: (2024)
par: Zheng, Dewu, et autres
Publié: (2024)
QuanBench: Benchmarking Quantum Code Generation with Large Language Models
par: Guo, Xiaoyu, et autres
Publié: (2025)
par: Guo, Xiaoyu, et autres
Publié: (2025)
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
par: Wang, Sizhe, et autres
Publié: (2025)
par: Wang, Sizhe, et autres
Publié: (2025)
OSS-Bench: Benchmark Generator for Coding LLMs
par: Jiang, Yuancheng, et autres
Publié: (2025)
par: Jiang, Yuancheng, et autres
Publié: (2025)
EDIT-Bench: Evaluating LLM Abilities to Perform Real-World Instructed Code Edits
par: Chi, Wayne, et autres
Publié: (2025)
par: Chi, Wayne, et autres
Publié: (2025)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
par: Aleithan, Reem, et autres
Publié: (2024)
par: Aleithan, Reem, et autres
Publié: (2024)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
par: Jing, Huihao, et autres
Publié: (2026)
par: Jing, Huihao, et autres
Publié: (2026)
SemOpt: LLM-Driven Code Optimization via Rule-Based Analysis
par: Zhao, Yuwei, et autres
Publié: (2025)
par: Zhao, Yuwei, et autres
Publié: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
par: Jiang, Hongchao, et autres
Publié: (2025)
par: Jiang, Hongchao, et autres
Publié: (2025)
GitEvo: Code Evolution Analysis for Git Repositories
par: Hora, Andre
Publié: (2026)
par: Hora, Andre
Publié: (2026)
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
par: Zhang, Shudan, et autres
Publié: (2024)
par: Zhang, Shudan, et autres
Publié: (2024)
Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
par: Zhang, Zhe, et autres
Publié: (2025)
par: Zhang, Zhe, et autres
Publié: (2025)
Condor: A Code Discriminator Integrating General Semantics with Code Details
par: Liang, Qingyuan, et autres
Publié: (2024)
par: Liang, Qingyuan, et autres
Publié: (2024)
RubberDuckBench: A Benchmark for AI Coding Assistants
par: Mohammed, Ferida, et autres
Publié: (2026)
par: Mohammed, Ferida, et autres
Publié: (2026)
EvoConfig: Self-Evolving Multi-Agent Systems for Efficient Autonomous Environment Configuration
par: Guo, Xinshuai, et autres
Publié: (2026)
par: Guo, Xinshuai, et autres
Publié: (2026)
Scaling Test-Driven Code Generation from Functions to Classes: An Empirical Study
par: Liang, Yunhao, et autres
Publié: (2026)
par: Liang, Yunhao, et autres
Publié: (2026)
Software Self-Extension with SelfEvolve: an Agentic Architecture for Runtime Code Generation
par: Fahim, Md Asif Iqbal, et autres
Publié: (2026)
par: Fahim, Md Asif Iqbal, et autres
Publié: (2026)
Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
par: Li, Zhuohao, et autres
Publié: (2025)
par: Li, Zhuohao, et autres
Publié: (2025)
Code vs Serialized AST Inputs for LLM-Based Code Summarization: An Empirical Study
par: Dong, Shijia, et autres
Publié: (2026)
par: Dong, Shijia, et autres
Publié: (2026)
DSCodeBench: A Realistic Benchmark for Data Science Code Generation
par: Ouyang, Shuyin, et autres
Publié: (2025)
par: Ouyang, Shuyin, et autres
Publié: (2025)
SimdBench: Benchmarking Large Language Models for SIMD-Intrinsic Code Generation
par: He, Yibo, et autres
Publié: (2025)
par: He, Yibo, et autres
Publié: (2025)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
par: Chou, Jason, et autres
Publié: (2025)
par: Chou, Jason, et autres
Publié: (2025)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
par: Xie, Yiqing, et autres
Publié: (2024)
par: Xie, Yiqing, et autres
Publié: (2024)
CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback
par: Sun, Qiushi, et autres
Publié: (2025)
par: Sun, Qiushi, et autres
Publié: (2025)
Benchmarking and Studying the LLM-based Code Review
par: Zeng, Zhengran, et autres
Publié: (2025)
par: Zeng, Zhengran, et autres
Publié: (2025)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
par: Duston, Titouan, et autres
Publié: (2025)
par: Duston, Titouan, et autres
Publié: (2025)
HumanEvalComm: Benchmarking the Communication Competence of Code Generation for LLMs and LLM Agent
par: Wu, Jie JW, et autres
Publié: (2024)
par: Wu, Jie JW, et autres
Publié: (2024)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding & Reasoning Capabilities of CodeLLMs
par: Manh, Dung Nguyen, et autres
Publié: (2024)
par: Manh, Dung Nguyen, et autres
Publié: (2024)
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
par: Guo, Hanyang, et autres
Publié: (2025)
par: Guo, Hanyang, et autres
Publié: (2025)
MigrationBench: Repository-Level Code Migration Benchmark from Java 8
par: Liu, Linbo, et autres
Publié: (2025)
par: Liu, Linbo, et autres
Publié: (2025)
Is LLM-Generated Code More Maintainable \& Reliable than Human-Written Code?
par: Molison, Alfred Santa, et autres
Publié: (2025)
par: Molison, Alfred Santa, et autres
Publié: (2025)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
par: Li, Jia, et autres
Publié: (2026)
par: Li, Jia, et autres
Publié: (2026)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
par: Huang, Dong, et autres
Publié: (2024)
par: Huang, Dong, et autres
Publié: (2024)
HyClone: Bridging LLM Understanding and Dynamic Execution for Semantic Code Clone Detection
par: Liang, Yunhao, et autres
Publié: (2025)
par: Liang, Yunhao, et autres
Publié: (2025)
R2Code: A Self-Reflective LLM Framework for Requirements-to-Code Traceability
par: Wang, Yifei, et autres
Publié: (2026)
par: Wang, Yifei, et autres
Publié: (2026)
Documents similaires
-
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
par: Li, Jia, et autres
Publié: (2024) -
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
par: Li, Jia, et autres
Publié: (2024) -
RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation
par: Liang, Linxi, et autres
Publié: (2025) -
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
par: Xia, Chunqiu Steven, et autres
Publié: (2024) -
HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation
par: Zheng, Dewu, et autres
Publié: (2024)