CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Yiqing, Xie, Alex, Sheth, Divyanshu, Liu, Pengfei, Fried, Daniel, Rose, Carolyn |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RepoST: Scalable Repository-Level Coding Environment Construction with Sandbox Testing
von: Xie, Yiqing, et al.
Veröffentlicht: (2025)
von: Xie, Yiqing, et al.
Veröffentlicht: (2025)
Data Augmentation for Code Translation with Comparable Corpora and Multiple References
von: Xie, Yiqing, et al.
Veröffentlicht: (2023)
von: Xie, Yiqing, et al.
Veröffentlicht: (2023)
CodeRAG-Bench: Can Retrieval Augment Code Generation?
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
MetaLint: Easy-to-Hard Generalization for Code Linting
von: Naik, Atharva, et al.
Veröffentlicht: (2025)
von: Naik, Atharva, et al.
Veröffentlicht: (2025)
Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
von: Xie, Yiqing, et al.
Veröffentlicht: (2026)
von: Xie, Yiqing, et al.
Veröffentlicht: (2026)
CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
von: Naik, Atharva, et al.
Veröffentlicht: (2024)
von: Naik, Atharva, et al.
Veröffentlicht: (2024)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
GenX: Mastering Code and Test Generation with Execution Feedback
von: Wang, Nan, et al.
Veröffentlicht: (2024)
von: Wang, Nan, et al.
Veröffentlicht: (2024)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
von: Chou, Jason, et al.
Veröffentlicht: (2025)
von: Chou, Jason, et al.
Veröffentlicht: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback
von: Peng, Yun, et al.
Veröffentlicht: (2024)
von: Peng, Yun, et al.
Veröffentlicht: (2024)
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
von: Wang, Sizhe, et al.
Veröffentlicht: (2025)
von: Wang, Sizhe, et al.
Veröffentlicht: (2025)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
von: Yan, Weixiang, et al.
Veröffentlicht: (2023)
von: Yan, Weixiang, et al.
Veröffentlicht: (2023)
An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generation
von: Gandhi, Shubham, et al.
Veröffentlicht: (2025)
von: Gandhi, Shubham, et al.
Veröffentlicht: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
von: Tian, Yuchen, et al.
Veröffentlicht: (2024)
von: Tian, Yuchen, et al.
Veröffentlicht: (2024)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
von: Zheng, Tianyu, et al.
Veröffentlicht: (2024)
von: Zheng, Tianyu, et al.
Veröffentlicht: (2024)
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
von: Zheng, Dewu, et al.
Veröffentlicht: (2024)
von: Zheng, Dewu, et al.
Veröffentlicht: (2024)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
ClassEval-Pro: A Cross-Domain Benchmark for Class-Level Code Generation
von: Chen, Yeheng, et al.
Veröffentlicht: (2026)
von: Chen, Yeheng, et al.
Veröffentlicht: (2026)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval
von: Wang, Jiexin, et al.
Veröffentlicht: (2024)
von: Wang, Jiexin, et al.
Veröffentlicht: (2024)
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
von: Zhang, Chenchen, et al.
Veröffentlicht: (2025)
von: Zhang, Chenchen, et al.
Veröffentlicht: (2025)
DuET: Dual Execution for Test Output Prediction with Generated Code and Pseudocode
von: Han, Hojae, et al.
Veröffentlicht: (2026)
von: Han, Hojae, et al.
Veröffentlicht: (2026)
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
von: Zhao, Songwen, et al.
Veröffentlicht: (2025)
von: Zhao, Songwen, et al.
Veröffentlicht: (2025)
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
Mercury: A Code Efficiency Benchmark for Code Large Language Models
von: Du, Mingzhe, et al.
Veröffentlicht: (2024)
von: Du, Mingzhe, et al.
Veröffentlicht: (2024)
BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2025)
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2025)
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
von: Yang, Jie, et al.
Veröffentlicht: (2026)
von: Yang, Jie, et al.
Veröffentlicht: (2026)
InstructCoder: Instruction Tuning Large Language Models for Code Editing
von: Li, Kaixin, et al.
Veröffentlicht: (2023)
von: Li, Kaixin, et al.
Veröffentlicht: (2023)
IndustryCode: A Benchmark for Industry Code Generation
von: Zeng, Puyu, et al.
Veröffentlicht: (2026)
von: Zeng, Puyu, et al.
Veröffentlicht: (2026)
Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation
von: Pysklo, Hubert M., et al.
Veröffentlicht: (2026)
von: Pysklo, Hubert M., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RepoST: Scalable Repository-Level Coding Environment Construction with Sandbox Testing
von: Xie, Yiqing, et al.
Veröffentlicht: (2025) -
Data Augmentation for Code Translation with Comparable Corpora and Multiple References
von: Xie, Yiqing, et al.
Veröffentlicht: (2023) -
CodeRAG-Bench: Can Retrieval Augment Code Generation?
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024) -
MetaLint: Easy-to-Hard Generalization for Code Linting
von: Naik, Atharva, et al.
Veröffentlicht: (2025) -
Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
von: Xie, Yiqing, et al.
Veröffentlicht: (2026)