Gespeichert in:
| Hauptverfasser: | Guan, Batu, Wu, Xiao, Yuan, Yuanyuan, Li, Shaohua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2503.06643 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mercury: A Code Efficiency Benchmark for Code Large Language Models
von: Du, Mingzhe, et al.
Veröffentlicht: (2024)
von: Du, Mingzhe, et al.
Veröffentlicht: (2024)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
von: Chou, Jason, et al.
Veröffentlicht: (2025)
von: Chou, Jason, et al.
Veröffentlicht: (2025)
Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026)
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026)
Iterative Refinement of Project-Level Code Context for Precise Code Generation with Compiler Feedback
von: Bi, Zhangqian, et al.
Veröffentlicht: (2024)
von: Bi, Zhangqian, et al.
Veröffentlicht: (2024)
A Code Comprehension Benchmark for Large Language Models for Code
von: Havare, Jayant, et al.
Veröffentlicht: (2025)
von: Havare, Jayant, et al.
Veröffentlicht: (2025)
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
von: Wu, Jiarong, et al.
Veröffentlicht: (2025)
von: Wu, Jiarong, et al.
Veröffentlicht: (2025)
What's Wrong with Your Code Generated by Large Language Models? An Extensive Study
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
von: Dou, Shihan, et al.
Veröffentlicht: (2024)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
von: Chen, Simin, et al.
Veröffentlicht: (2025)
von: Chen, Simin, et al.
Veröffentlicht: (2025)
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
von: Jiang, Nan, et al.
Veröffentlicht: (2024)
von: Jiang, Nan, et al.
Veröffentlicht: (2024)
R2C2-Coder: Enhancing and Benchmarking Real-world Repository-level Code Completion Abilities of Code Large Language Models
von: Deng, Ken, et al.
Veröffentlicht: (2024)
von: Deng, Ken, et al.
Veröffentlicht: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
von: Huang, Yue, et al.
Veröffentlicht: (2023)
von: Huang, Yue, et al.
Veröffentlicht: (2023)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
von: Zhao, Songwen, et al.
Veröffentlicht: (2025)
von: Zhao, Songwen, et al.
Veröffentlicht: (2025)
Benchmarking Failures in Tool-Augmented Language Models
von: Treviño, Eduardo, et al.
Veröffentlicht: (2025)
von: Treviño, Eduardo, et al.
Veröffentlicht: (2025)
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
von: Wang, Sizhe, et al.
Veröffentlicht: (2025)
von: Wang, Sizhe, et al.
Veröffentlicht: (2025)
Large Language Models are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence Tasks
von: Yang, Kang, et al.
Veröffentlicht: (2025)
von: Yang, Kang, et al.
Veröffentlicht: (2025)
Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval
von: Wang, Jiexin, et al.
Veröffentlicht: (2024)
von: Wang, Jiexin, et al.
Veröffentlicht: (2024)
HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
von: Peng, Qiwei, et al.
Veröffentlicht: (2024)
von: Peng, Qiwei, et al.
Veröffentlicht: (2024)
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
von: Xie, Yiqing, et al.
Veröffentlicht: (2024)
von: Xie, Yiqing, et al.
Veröffentlicht: (2024)
Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
Uncertainty Awareness of Large Language Models Under Code Distribution Shifts: A Benchmark Study
von: Li, Yufei, et al.
Veröffentlicht: (2024)
von: Li, Yufei, et al.
Veröffentlicht: (2024)
CODEMENV: Benchmarking Large Language Models on Code Migration
von: Cheng, Keyuan, et al.
Veröffentlicht: (2025)
von: Cheng, Keyuan, et al.
Veröffentlicht: (2025)
CodeUpdateArena: Benchmarking Knowledge Editing on API Updates
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2024)
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2024)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task
von: Liu, Jingyao, et al.
Veröffentlicht: (2025)
von: Liu, Jingyao, et al.
Veröffentlicht: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
IndustryCode: A Benchmark for Industry Code Generation
von: Zeng, Puyu, et al.
Veröffentlicht: (2026)
von: Zeng, Puyu, et al.
Veröffentlicht: (2026)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
von: Zhang, William, et al.
Veröffentlicht: (2024)
von: Zhang, William, et al.
Veröffentlicht: (2024)
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
von: Liu, Kaiyuan, et al.
Veröffentlicht: (2025)
von: Liu, Kaiyuan, et al.
Veröffentlicht: (2025)
DI-BENCH: Benchmarking Large Language Models on Dependency Inference with Testable Repositories at Scale
von: Zhang, Linghao, et al.
Veröffentlicht: (2025)
von: Zhang, Linghao, et al.
Veröffentlicht: (2025)
MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems
von: Li, Kaixin, et al.
Veröffentlicht: (2024)
von: Li, Kaixin, et al.
Veröffentlicht: (2024)
ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
von: Ding, Xianzhong, et al.
Veröffentlicht: (2026)
von: Ding, Xianzhong, et al.
Veröffentlicht: (2026)
Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
von: Li, Zhuohao, et al.
Veröffentlicht: (2025)
von: Li, Zhuohao, et al.
Veröffentlicht: (2025)
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
von: Lam, Man Ho, et al.
Veröffentlicht: (2026)
von: Lam, Man Ho, et al.
Veröffentlicht: (2026)
ClassEval-Pro: A Cross-Domain Benchmark for Class-Level Code Generation
von: Chen, Yeheng, et al.
Veröffentlicht: (2026)
von: Chen, Yeheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Mercury: A Code Efficiency Benchmark for Code Large Language Models
von: Du, Mingzhe, et al.
Veröffentlicht: (2024) -
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
von: Chou, Jason, et al.
Veröffentlicht: (2025) -
Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026) -
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
von: Chen, Zaoyu, et al.
Veröffentlicht: (2026) -
Iterative Refinement of Project-Level Code Context for Precise Code Generation with Compiler Feedback
von: Bi, Zhangqian, et al.
Veröffentlicht: (2024)