CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pereira, Kristen, Sinha, Neelabh, Ghosh, Rajat, Dutta, Debojyoti |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Multi-Agent Framework for Stateful Inference-Time Search
von: Lalan, Arshika, et al.
Veröffentlicht: (2025)
von: Lalan, Arshika, et al.
Veröffentlicht: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024)
von: Li, Jia, et al.
Veröffentlicht: (2024)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
von: Yang, Jie, et al.
Veröffentlicht: (2026)
von: Yang, Jie, et al.
Veröffentlicht: (2026)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
von: Wang, Yanli, et al.
Veröffentlicht: (2024)
von: Wang, Yanli, et al.
Veröffentlicht: (2024)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
von: Zhao, Bingchen, et al.
Veröffentlicht: (2026)
von: Zhao, Bingchen, et al.
Veröffentlicht: (2026)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
von: Sonwane, Atharv, et al.
Veröffentlicht: (2026)
von: Sonwane, Atharv, et al.
Veröffentlicht: (2026)
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
von: Orlanski, Gabriel, et al.
Veröffentlicht: (2026)
UA-Code-Bench: A Competitive Programming Benchmark for Evaluating LLM Code Generation in Ukrainian
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases
von: Wong, Sherman, et al.
Veröffentlicht: (2025)
von: Wong, Sherman, et al.
Veröffentlicht: (2025)
FeatBench: Towards More Realistic Evaluation of Feature-level Code Generation
von: Chen, Haorui, et al.
Veröffentlicht: (2025)
von: Chen, Haorui, et al.
Veröffentlicht: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++?
von: Bhargava, Vaishnavi, et al.
Veröffentlicht: (2024)
von: Bhargava, Vaishnavi, et al.
Veröffentlicht: (2024)
CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
von: Naik, Atharva, et al.
Veröffentlicht: (2024)
von: Naik, Atharva, et al.
Veröffentlicht: (2024)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
von: Zhang, Zehua, et al.
Veröffentlicht: (2025)
von: Zhang, Zehua, et al.
Veröffentlicht: (2025)
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2024)
MOSS: Enabling Code-Driven Evolution and Context Management for AI Agents
von: Zhu, Ming, et al.
Veröffentlicht: (2024)
von: Zhu, Ming, et al.
Veröffentlicht: (2024)
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
von: Chen, Junkai, et al.
Veröffentlicht: (2025)
von: Chen, Junkai, et al.
Veröffentlicht: (2025)
CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
von: Guo, Jiawei, et al.
Veröffentlicht: (2024)
von: Guo, Jiawei, et al.
Veröffentlicht: (2024)
Collu-Bench: A Benchmark for Predicting Language Model Hallucinations in Code
von: Jiang, Nan, et al.
Veröffentlicht: (2024)
von: Jiang, Nan, et al.
Veröffentlicht: (2024)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
von: Duston, Titouan, et al.
Veröffentlicht: (2025)
Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development
von: Tran, Hung, et al.
Veröffentlicht: (2026)
von: Tran, Hung, et al.
Veröffentlicht: (2026)
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
von: Diddee, Harshita, et al.
Veröffentlicht: (2026)
von: Diddee, Harshita, et al.
Veröffentlicht: (2026)
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
von: Zheng, Dewu, et al.
Veröffentlicht: (2024)
von: Zheng, Dewu, et al.
Veröffentlicht: (2024)
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments
von: Han, Hojae, et al.
Veröffentlicht: (2025)
von: Han, Hojae, et al.
Veröffentlicht: (2025)
DebugBench: Evaluating Debugging Capability of Large Language Models
von: Tian, Runchu, et al.
Veröffentlicht: (2024)
von: Tian, Runchu, et al.
Veröffentlicht: (2024)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
von: Tu, Xinming, et al.
Veröffentlicht: (2026)
LocAgent: Graph-Guided LLM Agents for Code Localization
von: Chen, Zhaoling, et al.
Veröffentlicht: (2025)
von: Chen, Zhaoling, et al.
Veröffentlicht: (2025)
CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
von: Sutawika, Lintang, et al.
Veröffentlicht: (2026)
von: Sutawika, Lintang, et al.
Veröffentlicht: (2026)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
von: Chi, Wayne, et al.
Veröffentlicht: (2026)
von: Chi, Wayne, et al.
Veröffentlicht: (2026)
Training Versatile Coding Agents in Synthetic Environments
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
Dissecting the SWE-Bench Leaderboards: Profiling Submitters and Architectures of LLM- and Agent-Based Repair Systems
von: Martinez, Matias, et al.
Veröffentlicht: (2025)
von: Martinez, Matias, et al.
Veröffentlicht: (2025)
AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
von: Trivedi, Harsh, et al.
Veröffentlicht: (2024)
von: Trivedi, Harsh, et al.
Veröffentlicht: (2024)
Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
Effective Harness Engineering for Algorithm Discovery with Coding Agents
von: Ishibashi, Yoichi, et al.
Veröffentlicht: (2026)
von: Ishibashi, Yoichi, et al.
Veröffentlicht: (2026)
MERA Code: A Unified Framework for Evaluating Code Generation Across Tasks
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities
von: Wang, Hanbin, et al.
Veröffentlicht: (2025)
von: Wang, Hanbin, et al.
Veröffentlicht: (2025)
RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
von: Gautam, Dhruv, et al.
Veröffentlicht: (2025)
von: Gautam, Dhruv, et al.
Veröffentlicht: (2025)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
von: Wang, Peiding, et al.
Veröffentlicht: (2025)
von: Wang, Peiding, et al.
Veröffentlicht: (2025)
Can Coding Agents Reproduce Findings in Computational Materials Science?
von: Huang, Ziyang, et al.
Veröffentlicht: (2026)
von: Huang, Ziyang, et al.
Veröffentlicht: (2026)
Asymmetric Goal Drift in Coding Agents Under Value Conflict
von: Saebo, Magnus, et al.
Veröffentlicht: (2026)
von: Saebo, Magnus, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Multi-Agent Framework for Stateful Inference-Time Search
von: Lalan, Arshika, et al.
Veröffentlicht: (2025) -
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
von: Li, Jia, et al.
Veröffentlicht: (2024) -
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
von: Yang, Jie, et al.
Veröffentlicht: (2026) -
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
von: Wang, Yanli, et al.
Veröffentlicht: (2024) -
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
von: Zhao, Bingchen, et al.
Veröffentlicht: (2026)