CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Jiang, Hongchao, Chen, Yiming, Cao, Yushi, Lee, Hung-yi, Tan, Robby T. |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Rethinking Code Refinement: Learning to Judge Code Efficiency
par: Seo, Minju, et autres
Publié: (2024)
par: Seo, Minju, et autres
Publié: (2024)
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
par: Jiang, Nan, et autres
Publié: (2024)
par: Jiang, Nan, et autres
Publié: (2024)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
par: Orlanski, Gabriel, et autres
Publié: (2026)
par: Orlanski, Gabriel, et autres
Publié: (2026)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
par: Syromiatnikov, Mykyta, et autres
Publié: (2025)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
par: Zhou, Xin, et autres
Publié: (2025)
par: Zhou, Xin, et autres
Publié: (2025)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
par: Moon, Jiwon, et autres
Publié: (2025)
par: Moon, Jiwon, et autres
Publié: (2025)
LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?
par: Zheng, Zihan, et autres
Publié: (2025)
par: Zheng, Zihan, et autres
Publié: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
par: Zhuo, Terry Yue, et autres
Publié: (2024)
par: Zhuo, Terry Yue, et autres
Publié: (2024)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
par: Lai, Peng, et autres
Publié: (2026)
par: Lai, Peng, et autres
Publié: (2026)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
par: Yang, Jie, et autres
Publié: (2026)
par: Yang, Jie, et autres
Publié: (2026)
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
par: Zheng, Dewu, et autres
Publié: (2024)
par: Zheng, Dewu, et autres
Publié: (2024)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
par: Wang, Yanli, et autres
Publié: (2024)
par: Wang, Yanli, et autres
Publié: (2024)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
par: Zhao, Bingchen, et autres
Publié: (2026)
par: Zhao, Bingchen, et autres
Publié: (2026)
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
par: Zhao, Yuwei, et autres
Publié: (2024)
par: Zhao, Yuwei, et autres
Publié: (2024)
Pragmatic Reasoning improves LLM Code Generation
par: Cao, Zhuchen, et autres
Publié: (2025)
par: Cao, Zhuchen, et autres
Publié: (2025)
IndustryCode: A Benchmark for Industry Code Generation
par: Zeng, Puyu, et autres
Publié: (2026)
par: Zeng, Puyu, et autres
Publié: (2026)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
par: Duston, Titouan, et autres
Publié: (2025)
par: Duston, Titouan, et autres
Publié: (2025)
A Code Comprehension Benchmark for Large Language Models for Code
par: Havare, Jayant, et autres
Publié: (2025)
par: Havare, Jayant, et autres
Publié: (2025)
FeatBench: Towards More Realistic Evaluation of Feature-level Code Generation
par: Chen, Haorui, et autres
Publié: (2025)
par: Chen, Haorui, et autres
Publié: (2025)
AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
par: Wang, Ruiqi, et autres
Publié: (2025)
par: Wang, Ruiqi, et autres
Publié: (2025)
Crystal: Illuminating LLM Abilities on Language and Code
par: Tao, Tianhua, et autres
Publié: (2024)
par: Tao, Tianhua, et autres
Publié: (2024)
MERA Code: A Unified Framework for Evaluating Code Generation Across Tasks
par: Chervyakov, Artem, et autres
Publié: (2025)
par: Chervyakov, Artem, et autres
Publié: (2025)
LocAgent: Graph-Guided LLM Agents for Code Localization
par: Chen, Zhaoling, et autres
Publié: (2025)
par: Chen, Zhaoling, et autres
Publié: (2025)
Issue Localization via LLM-Driven Iterative Code Graph Searching
par: Jiang, Zhonghao, et autres
Publié: (2025)
par: Jiang, Zhonghao, et autres
Publié: (2025)
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments
par: Han, Hojae, et autres
Publié: (2025)
par: Han, Hojae, et autres
Publié: (2025)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
par: Zhang, William, et autres
Publié: (2024)
par: Zhang, William, et autres
Publié: (2024)
Transducer Tuning: Efficient Model Adaptation for Software Tasks Using Code Property Graphs
par: Yusuf, Imam Nur Bani, et autres
Publié: (2024)
par: Yusuf, Imam Nur Bani, et autres
Publié: (2024)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
par: Tu, Xinming, et autres
Publié: (2026)
par: Tu, Xinming, et autres
Publié: (2026)
Advancing Language Models for Code-related Tasks
par: Tian, Zhao
Publié: (2026)
par: Tian, Zhao
Publié: (2026)
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
par: Yan, Weixiang, et autres
Publié: (2023)
par: Yan, Weixiang, et autres
Publié: (2023)
Stingy Context: 18:1 Hierarchical Code Compression for LLM Auto-Coding
par: Ostby, David Linus
Publié: (2026)
par: Ostby, David Linus
Publié: (2026)
Verification Limits Code LLM Training
par: Gureja, Srishti, et autres
Publié: (2025)
par: Gureja, Srishti, et autres
Publié: (2025)
Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models
par: Zheng, Jiasheng, et autres
Publié: (2024)
par: Zheng, Jiasheng, et autres
Publié: (2024)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
par: Sonwane, Atharv, et autres
Publié: (2026)
par: Sonwane, Atharv, et autres
Publié: (2026)
CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
par: Pereira, Kristen, et autres
Publié: (2026)
par: Pereira, Kristen, et autres
Publié: (2026)
CodeR: Issue Resolving with Multi-Agent and Task Graphs
par: Chen, Dong, et autres
Publié: (2024)
par: Chen, Dong, et autres
Publié: (2024)
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
par: Zhang, Ziyao, et autres
Publié: (2024)
par: Zhang, Ziyao, et autres
Publié: (2024)
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
par: Jiang, Xue, et autres
Publié: (2025)
par: Jiang, Xue, et autres
Publié: (2025)
CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
par: Guo, Jiawei, et autres
Publié: (2024)
par: Guo, Jiawei, et autres
Publié: (2024)
Documents similaires
-
Rethinking Code Refinement: Learning to Judge Code Efficiency
par: Seo, Minju, et autres
Publié: (2024) -
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
par: Jiang, Nan, et autres
Publié: (2024) -
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
par: Orlanski, Gabriel, et autres
Publié: (2026) -
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
par: Syromiatnikov, Mykyta, et autres
Publié: (2025) -
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
par: Zhou, Xin, et autres
Publié: (2025)