Benchmarking and Studying the LLM-based Code Review
Fuente:
arXiv
Saved in:
| Main Authors: | Zeng, Zhengran, Shi, Ruikai, Han, Keke, Li, Yixin, Sun, Kaicheng, Wang, Yidong, Yu, Zhuohao, Xie, Rui, Ye, Wei, Zhang, Shikun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development
by: Zeng, Zhengran, et al.
Published: (2025)
by: Zeng, Zhengran, et al.
Published: (2025)
CodeShell Technical Report
by: Xie, Rui, et al.
Published: (2024)
by: Xie, Rui, et al.
Published: (2024)
An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models
by: Xing, Chengli, et al.
Published: (2026)
by: Xing, Chengli, et al.
Published: (2026)
CoderUJB: An Executable and Unified Java Benchmark for Practical Programming Scenarios
by: Zeng, Zhengran, et al.
Published: (2024)
by: Zeng, Zhengran, et al.
Published: (2024)
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
by: Yu, Zhuohao, et al.
Published: (2024)
by: Yu, Zhuohao, et al.
Published: (2024)
Seed&Steer: Guiding Large Language Models with Compilable Prefix and Branch Signals for Unit Test Generation
by: Zhou, Shuaiyu, et al.
Published: (2025)
by: Zhou, Shuaiyu, et al.
Published: (2025)
ISC4DGF: Enhancing Directed Grey-box Fuzzing with LLM-Driven Initial Seed Corpus Generation
by: Xu, Yijiang, et al.
Published: (2024)
by: Xu, Yijiang, et al.
Published: (2024)
A Survey on Evaluating Large Language Models in Code Generation Tasks
by: Chen, Liguo, et al.
Published: (2024)
by: Chen, Liguo, et al.
Published: (2024)
Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
by: Li, Zhuohao, et al.
Published: (2025)
by: Li, Zhuohao, et al.
Published: (2025)
HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation
by: Zheng, Dewu, et al.
Published: (2024)
by: Zheng, Dewu, et al.
Published: (2024)
CLARC: C/C++ Benchmark for Robust Code Search
by: Wang, Kaicheng, et al.
Published: (2026)
by: Wang, Kaicheng, et al.
Published: (2026)
Inducing Vulnerable Code Generation in LLM Coding Assistants
by: Zeng, Binqi, et al.
Published: (2025)
by: Zeng, Binqi, et al.
Published: (2025)
A Survey of Code Review Benchmarks and Evaluation Practices in Pre-LLM and LLM Era
by: Khan, Taufiqul Islam, et al.
Published: (2026)
by: Khan, Taufiqul Islam, et al.
Published: (2026)
Evaluating LLM-Generated Code: A Benchmark and Developer Study
by: Szych, Joanna, et al.
Published: (2026)
by: Szych, Joanna, et al.
Published: (2026)
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
by: Liu, Zhuoyao, et al.
Published: (2026)
by: Liu, Zhuoyao, et al.
Published: (2026)
SAINT: Service-level Integration Test Generation with Program Analysis and LLM-based Agents
by: Pan, Rangeet, et al.
Published: (2025)
by: Pan, Rangeet, et al.
Published: (2025)
Rethinking Code Review Workflows with LLM Assistance: An Empirical Study
by: Aðalsteinsson, Fannar Steinn, et al.
Published: (2025)
by: Aðalsteinsson, Fannar Steinn, et al.
Published: (2025)
BitsAI-CR: Automated Code Review via LLM in Practice
by: Sun, Tao, et al.
Published: (2025)
by: Sun, Tao, et al.
Published: (2025)
RLCoder: Reinforcement Learning for Repository-Level Code Completion
by: Wang, Yanlin, et al.
Published: (2024)
by: Wang, Yanlin, et al.
Published: (2024)
When to Stop? Towards Efficient Code Generation in LLMs with Excess Token Prevention
by: Guo, Lianghong, et al.
Published: (2024)
by: Guo, Lianghong, et al.
Published: (2024)
An Empirical Study of Interaction Smells in Multi-Turn Human-LLM Collaborative Code Generation
by: Zhang, Binquan, et al.
Published: (2026)
by: Zhang, Binquan, et al.
Published: (2026)
Environment-in-the-Loop: Rethinking Code Migration with LLM-based Agents
by: Li, Xiang, et al.
Published: (2026)
by: Li, Xiang, et al.
Published: (2026)
Code Review Automation Via Multi-task Federated LLM -- An Empirical Study
by: Kumar, Jahnavi, et al.
Published: (2024)
by: Kumar, Jahnavi, et al.
Published: (2024)
SGCR: A Specification-Grounded Framework for Trustworthy LLM Code Review
by: Wang, Kai, et al.
Published: (2025)
by: Wang, Kai, et al.
Published: (2025)
Prompt Engineering for Requirements Engineering: A Literature Review and Roadmap
by: Huang, Kaicheng, et al.
Published: (2025)
by: Huang, Kaicheng, et al.
Published: (2025)
LAURA: Enhancing Code Review Generation with Context-Enriched Retrieval-Augmented LLM
by: Zhang, Yuxin, et al.
Published: (2025)
by: Zhang, Yuxin, et al.
Published: (2025)
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
by: Xie, Yiqing, et al.
Published: (2024)
by: Xie, Yiqing, et al.
Published: (2024)
ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS
by: Xie, Bang, et al.
Published: (2026)
by: Xie, Bang, et al.
Published: (2026)
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
RedCode: Risky Code Execution and Generation Benchmark for Code Agents
by: Guo, Chengquan, et al.
Published: (2024)
by: Guo, Chengquan, et al.
Published: (2024)
Improving LLM-Based Go Code Review through Issue-List Generation and Context Augmentation
by: Sun, Kexin, et al.
Published: (2026)
by: Sun, Kexin, et al.
Published: (2026)
CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
by: Guo, Hanyang, et al.
Published: (2025)
by: Guo, Hanyang, et al.
Published: (2025)
From UI to Code: Mobile Ads Detection via LLM-Unified Static-Dynamic Analysis
by: Ma, Shang, et al.
Published: (2026)
by: Ma, Shang, et al.
Published: (2026)
On the Effectiveness of Training Data Optimization for LLM-based Code Generation: An Empirical Study
by: Kuang, Shiqi, et al.
Published: (2025)
by: Kuang, Shiqi, et al.
Published: (2025)
Context-Aware CodeLLM Eviction for AI-assisted Coding
by: Thangarajah, Kishanthan, et al.
Published: (2025)
by: Thangarajah, Kishanthan, et al.
Published: (2025)
ClassEval-Pro: A Cross-Domain Benchmark for Class-Level Code Generation
by: Chen, Yeheng, et al.
Published: (2026)
by: Chen, Yeheng, et al.
Published: (2026)
Code Review Agent Benchmark
by: Zhang, Yuntong, et al.
Published: (2026)
by: Zhang, Yuntong, et al.
Published: (2026)
Requirements-Driven Automated Software Testing: A Systematic Review
by: Wang, Fanyu, et al.
Published: (2025)
by: Wang, Fanyu, et al.
Published: (2025)
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation
by: Zhang, Binquan, et al.
Published: (2025)
by: Zhang, Binquan, et al.
Published: (2025)
IntrinTrans: LLM-based Intrinsic Code Translator for RISC-V Vector
by: Han, Liutong, et al.
Published: (2025)
by: Han, Liutong, et al.
Published: (2025)
Similar Items
-
Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development
by: Zeng, Zhengran, et al.
Published: (2025) -
CodeShell Technical Report
by: Xie, Rui, et al.
Published: (2024) -
An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models
by: Xing, Chengli, et al.
Published: (2026) -
CoderUJB: An Executable and Unified Java Benchmark for Practical Programming Scenarios
by: Zeng, Zhengran, et al.
Published: (2024) -
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
by: Yu, Zhuohao, et al.
Published: (2024)