Benchmarking LLMs for Fine-Grained Code Review with Enriched Context in Practice
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Ruida, Wang, Xinchen, Wen, Xin-Cheng, Zhang, Zhao, Jiang, Bo, Gao, Pengfei, Peng, Chao, Gao, Cuiyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models
von: Hu, Ruida, et al.
Veröffentlicht: (2024)
von: Hu, Ruida, et al.
Veröffentlicht: (2024)
CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation
von: Wang, Xinchen, et al.
Veröffentlicht: (2025)
von: Wang, Xinchen, et al.
Veröffentlicht: (2025)
Evaluating Repository-level Software Documentation via Question Answering and Feature-Driven Development
von: Wang, Xinchen, et al.
Veröffentlicht: (2026)
von: Wang, Xinchen, et al.
Veröffentlicht: (2026)
CodeRepoQA: A Large-scale Benchmark for Software Engineering Question Answering
von: Hu, Ruida, et al.
Veröffentlicht: (2024)
von: Hu, Ruida, et al.
Veröffentlicht: (2024)
SR-Eval: Evaluating LLMs on Code Generation under Stepwise Requirement Refinement
von: Zhan, Zexun, et al.
Veröffentlicht: (2025)
von: Zhan, Zexun, et al.
Veröffentlicht: (2025)
Repo2Run: Automated Building Executable Environment for Code Repository at Scale
von: Hu, Ruida, et al.
Veröffentlicht: (2025)
von: Hu, Ruida, et al.
Veröffentlicht: (2025)
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
von: Hu, Ruida, et al.
Veröffentlicht: (2026)
von: Hu, Ruida, et al.
Veröffentlicht: (2026)
AEGIS: An Agent-based Framework for General Bug Reproduction from Issue Descriptions
von: Wang, Xinchen, et al.
Veröffentlicht: (2024)
von: Wang, Xinchen, et al.
Veröffentlicht: (2024)
ReposVul: A Repository-Level High-Quality Vulnerability Dataset
von: Wang, Xinchen, et al.
Veröffentlicht: (2024)
von: Wang, Xinchen, et al.
Veröffentlicht: (2024)
VulEval: Towards Repository-Level Evaluation of Software Vulnerability Detection
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2024)
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2024)
What Makes Good In-context Demonstrations for Code Intelligence Tasks with LLMs?
von: Gao, Shuzheng, et al.
Veröffentlicht: (2023)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2023)
RepoMasterEval: Evaluating Code Completion via Real-World Repositories
von: Wu, Qinyun, et al.
Veröffentlicht: (2024)
von: Wu, Qinyun, et al.
Veröffentlicht: (2024)
AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex Code
von: Feng, Jia, et al.
Veröffentlicht: (2024)
von: Feng, Jia, et al.
Veröffentlicht: (2024)
Less is More? An Empirical Study on Configuration Issues in Python PyPI Ecosystem
von: Peng, Yun, et al.
Veröffentlicht: (2023)
von: Peng, Yun, et al.
Veröffentlicht: (2023)
Game Rewards Vulnerabilities: Software Vulnerability Detection with Zero-Sum Game and Prototype Learning
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2024)
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2024)
RAG or Fine-tuning? A Comparative Study on LCMs-based Code Completion in Industry
von: Wang, Chaozheng, et al.
Veröffentlicht: (2025)
von: Wang, Chaozheng, et al.
Veröffentlicht: (2025)
Search-Based LLMs for Code Optimization
von: Gao, Shuzheng, et al.
Veröffentlicht: (2024)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2024)
Towards Mitigating API Hallucination in Code Generated by LLMs with Hierarchical Dependency Aware
von: Chen, Yujia, et al.
Veröffentlicht: (2025)
von: Chen, Yujia, et al.
Veröffentlicht: (2025)
MarsCode Agent: AI-native Automated Bug Fixing
von: Liu, Yizhou, et al.
Veröffentlicht: (2024)
von: Liu, Yizhou, et al.
Veröffentlicht: (2024)
Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling
von: Trae Research Team, et al.
Veröffentlicht: (2025)
von: Trae Research Team, et al.
Veröffentlicht: (2025)
A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
von: Gao, Cuiyun, et al.
Veröffentlicht: (2025)
von: Gao, Cuiyun, et al.
Veröffentlicht: (2025)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
von: Wu, Fan, et al.
Veröffentlicht: (2026)
von: Wu, Fan, et al.
Veröffentlicht: (2026)
A Roadmap on Modern Code Review: Challenges and Opportunities
von: Yang, Zezhou, et al.
Veröffentlicht: (2024)
von: Yang, Zezhou, et al.
Veröffentlicht: (2024)
Boosting Vulnerability Detection of LLMs via Curriculum Preference Optimization with Synthetic Reasoning Data
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2025)
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2025)
CoRE: A Fine-Grained Code Reasoning Benchmark Beyond Output Prediction
von: Gao, Jun, et al.
Veröffentlicht: (2026)
von: Gao, Jun, et al.
Veröffentlicht: (2026)
LAURA: Enhancing Code Review Generation with Context-Enriched Retrieval-Augmented LLM
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025)
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025)
MLLM-Based UI2Code Automation Guided by UI Layout Information
von: Wu, Fan, et al.
Veröffentlicht: (2025)
von: Wu, Fan, et al.
Veröffentlicht: (2025)
Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code Models
von: Gao, Shuzheng, et al.
Veröffentlicht: (2024)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2024)
The Current Challenges of Software Engineering in the Era of Large Language Models
von: Gao, Cuiyun, et al.
Veröffentlicht: (2024)
von: Gao, Cuiyun, et al.
Veröffentlicht: (2024)
SEER: Enhancing Chain-of-Thought Code Generation through Self-Exploring Deep Reasoning
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
Deep Learning Based Code Generation Methods: Literature Review
von: Yang, Zezhou, et al.
Veröffentlicht: (2023)
von: Yang, Zezhou, et al.
Veröffentlicht: (2023)
A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback
von: Duan, Guoliang, et al.
Veröffentlicht: (2025)
von: Duan, Guoliang, et al.
Veröffentlicht: (2025)
More with Less: An Empirical Study of Turn-Control Strategies for Efficient Coding Agents
von: Gao, Pengfei, et al.
Veröffentlicht: (2025)
von: Gao, Pengfei, et al.
Veröffentlicht: (2025)
An Empirical Study of Retrieval-Augmented Code Generation: Challenges and Opportunities
von: Yang, Zezhou, et al.
Veröffentlicht: (2025)
von: Yang, Zezhou, et al.
Veröffentlicht: (2025)
Schedule-and-Calibrate: Utility-Guided Multi-Task Reinforcement Learning for Code LLMs
von: Chen, Yujia, et al.
Veröffentlicht: (2026)
von: Chen, Yujia, et al.
Veröffentlicht: (2026)
SCALE: Constructing Structured Natural Language Comment Trees for Software Vulnerability Detection
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2024)
von: Wen, Xin-Cheng, et al.
Veröffentlicht: (2024)
A Deep Dive into Retrieval-Augmented Generation for Code Completion: Experience on WeChat
von: Yang, Zezhou, et al.
Veröffentlicht: (2025)
von: Yang, Zezhou, et al.
Veröffentlicht: (2025)
An Empirical Study of Knowledge Distillation for Code Understanding Tasks
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
A Closer Look into Transformer-Based Code Intelligence Through Code Transformation: Challenges and Opportunities
von: Li, Yaoxian, et al.
Veröffentlicht: (2022)
von: Li, Yaoxian, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models
von: Hu, Ruida, et al.
Veröffentlicht: (2024) -
CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation
von: Wang, Xinchen, et al.
Veröffentlicht: (2025) -
Evaluating Repository-level Software Documentation via Question Answering and Feature-Driven Development
von: Wang, Xinchen, et al.
Veröffentlicht: (2026) -
CodeRepoQA: A Large-scale Benchmark for Software Engineering Question Answering
von: Hu, Ruida, et al.
Veröffentlicht: (2024) -
SR-Eval: Evaluating LLMs on Code Generation under Stepwise Requirement Refinement
von: Zhan, Zexun, et al.
Veröffentlicht: (2025)