EffiBench: Benchmarking the Efficiency of Automatically Generated Code
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Dong, Qing, Yuhao, Shang, Weiyi, Cui, Heming, Zhang, Jie M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EffiLearner: Enhancing Efficiency of Generated Code via Self-Optimization
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated Code
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
EffiSkill: Agent Skill Based Automated Code Efficiency Optimization
von: Wang, Zimu, et al.
Veröffentlicht: (2026)
von: Wang, Zimu, et al.
Veröffentlicht: (2026)
Measuring the Influence of Incorrect Code on Test Generation
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
von: Chou, Jason, et al.
Veröffentlicht: (2025)
von: Chou, Jason, et al.
Veröffentlicht: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation
von: Huang, Dong, et al.
Veröffentlicht: (2023)
von: Huang, Dong, et al.
Veröffentlicht: (2023)
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
von: Wang, Sizhe, et al.
Veröffentlicht: (2025)
von: Wang, Sizhe, et al.
Veröffentlicht: (2025)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
von: Xie, Yiqing, et al.
Veröffentlicht: (2024)
von: Xie, Yiqing, et al.
Veröffentlicht: (2024)
EffiPair: Improving the Efficiency of LLM-generated Code with Relative Contrastive Feedback
von: Hajizadeh, Samira, et al.
Veröffentlicht: (2026)
von: Hajizadeh, Samira, et al.
Veröffentlicht: (2026)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
von: Zhang, Chenchen, et al.
Veröffentlicht: (2025)
von: Zhang, Chenchen, et al.
Veröffentlicht: (2025)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
Mercury: A Code Efficiency Benchmark for Code Large Language Models
von: Du, Mingzhe, et al.
Veröffentlicht: (2024)
von: Du, Mingzhe, et al.
Veröffentlicht: (2024)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
CodeRAG-Bench: Can Retrieval Augment Code Generation?
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
von: Wang, Zora Zhiruo, et al.
Veröffentlicht: (2024)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
von: Huang, Dong, et al.
Veröffentlicht: (2025)
von: Huang, Dong, et al.
Veröffentlicht: (2025)
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
CommitBench: A Benchmark for Commit Message Generation
von: Schall, Maximilian, et al.
Veröffentlicht: (2024)
von: Schall, Maximilian, et al.
Veröffentlicht: (2024)
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
von: Zheng, Dewu, et al.
Veröffentlicht: (2024)
von: Zheng, Dewu, et al.
Veröffentlicht: (2024)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
von: Yang, Jie, et al.
Veröffentlicht: (2026)
von: Yang, Jie, et al.
Veröffentlicht: (2026)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
The Role of DevOps in Enhancing Enterprise Software Delivery Success through R&D Efficiency and Source Code Management
von: Cui, Jun
Veröffentlicht: (2024)
von: Cui, Jun
Veröffentlicht: (2024)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts
von: Zhang, Shudan, et al.
Veröffentlicht: (2024)
von: Zhang, Shudan, et al.
Veröffentlicht: (2024)
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
von: Jiang, Nan, et al.
Veröffentlicht: (2024)
von: Jiang, Nan, et al.
Veröffentlicht: (2024)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
von: Zhao, Songwen, et al.
Veröffentlicht: (2025)
von: Zhao, Songwen, et al.
Veröffentlicht: (2025)
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
DSCodeBench: A Realistic Benchmark for Data Science Code Generation
von: Ouyang, Shuyin, et al.
Veröffentlicht: (2025)
von: Ouyang, Shuyin, et al.
Veröffentlicht: (2025)
CONCUR: Benchmarking LLMs for Concurrent Code Generation
von: Huang, Jue, et al.
Veröffentlicht: (2026)
von: Huang, Jue, et al.
Veröffentlicht: (2026)
ClassEval-Pro: A Cross-Domain Benchmark for Class-Level Code Generation
von: Chen, Yeheng, et al.
Veröffentlicht: (2026)
von: Chen, Yeheng, et al.
Veröffentlicht: (2026)
StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
von: Lin, Jiahang, et al.
Veröffentlicht: (2026)
von: Lin, Jiahang, et al.
Veröffentlicht: (2026)
IndustryCode: A Benchmark for Industry Code Generation
von: Zeng, Puyu, et al.
Veröffentlicht: (2026)
von: Zeng, Puyu, et al.
Veröffentlicht: (2026)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
von: Wang, Yanli, et al.
Veröffentlicht: (2024)
von: Wang, Yanli, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EffiLearner: Enhancing Efficiency of Generated Code via Self-Optimization
von: Huang, Dong, et al.
Veröffentlicht: (2024) -
EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning
von: Huang, Dong, et al.
Veröffentlicht: (2024) -
EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated Code
von: Qing, Yuhao, et al.
Veröffentlicht: (2025) -
EffiSkill: Agent Skill Based Automated Code Efficiency Optimization
von: Wang, Zimu, et al.
Veröffentlicht: (2026) -
Measuring the Influence of Incorrect Code on Test Generation
von: Huang, Dong, et al.
Veröffentlicht: (2024)