A Survey of Code Review Benchmarks and Evaluation Practices in Pre-LLM and LLM Era
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Khan, Taufiqul Islam, Wang, Shaowei, Zhang, Haoxiang, Chen, Tse-Hsun |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
CODEPROMPTZIP: Code-specific Prompt Compression for Retrieval-Augmented Generation in Coding Tasks with LMs
par: He, Pengfei, et autres
Publié: (2025)
par: He, Pengfei, et autres
Publié: (2025)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
par: Xu, Yisen, et autres
Publié: (2026)
par: Xu, Yisen, et autres
Publié: (2026)
Evaluating the Effectiveness and Efficiency of Demonstration Retrievers in RAG for Coding Tasks
par: He, Pengfei, et autres
Publié: (2024)
par: He, Pengfei, et autres
Publié: (2024)
SLICET5: Static Program Slicing using Language Models with Copy Mechanism and Constrained Decoding
par: He, Pengfei, et autres
Publié: (2025)
par: He, Pengfei, et autres
Publié: (2025)
Identifying Performance-Sensitive Configurations in Software Systems through Code Analysis with LLM Agents
par: Wang, Zehao, et autres
Publié: (2024)
par: Wang, Zehao, et autres
Publié: (2024)
Studying the Impact of Early Test Termination Due to Assertion Failure on Code Coverage and Spectrum-based Fault Localization
par: Uddin, Md. Ashraf, et autres
Publié: (2025)
par: Uddin, Md. Ashraf, et autres
Publié: (2025)
Evaluating Software Process Models for Multi-Agent Class-Level Code Generation
par: Shafin, Wasique Islam, et autres
Publié: (2025)
par: Shafin, Wasique Islam, et autres
Publié: (2025)
When LLMs Lag Behind: Knowledge Conflicts from Evolving APIs in Code Generation
par: Ashik, Ahmed Nusayer, et autres
Publié: (2026)
par: Ashik, Ahmed Nusayer, et autres
Publié: (2026)
Towards Better Graph Neural Network-based Fault Localization Through Enhanced Code Representation
par: Rafi, Md Nakhla, et autres
Publié: (2024)
par: Rafi, Md Nakhla, et autres
Publié: (2024)
ZS4C: Zero-Shot Synthesis of Compilable Code for Incomplete Code Snippets using LLMs
par: Kabir, Azmain, et autres
Publié: (2024)
par: Kabir, Azmain, et autres
Publié: (2024)
LLM-Based Detection of Tangled Code Changes for Higher-Quality Method-Level Bug Datasets
par: Opu, Md Nahidul Islam, et autres
Publié: (2025)
par: Opu, Md Nahidul Islam, et autres
Publié: (2025)
A Comprehensive Framework for Evaluating API-oriented Code Generation in Large Language Models
par: Wu, Yixi, et autres
Publié: (2024)
par: Wu, Yixi, et autres
Publié: (2024)
A Multi-Agent Approach to Fault Localization via Graph-Based Retrieval and Reflexion
par: Rafi, Md Nakhla, et autres
Publié: (2024)
par: Rafi, Md Nakhla, et autres
Publié: (2024)
Studying and Recommending Information Highlighting in Stack Overflow Answers
par: Ahmed, Shahla Shaan, et autres
Publié: (2024)
par: Ahmed, Shahla Shaan, et autres
Publié: (2024)
Benchmarking and Studying the LLM-based Code Review
par: Zeng, Zhengran, et autres
Publié: (2025)
par: Zeng, Zhengran, et autres
Publié: (2025)
Discovery of Timeline and Crowd Reaction of Software Vulnerability Disclosures
par: Heng, Yi Wen, et autres
Publié: (2024)
par: Heng, Yi Wen, et autres
Publié: (2024)
Compressing Code Context for LLM-based Issue Resolution
par: Jia, Haoxiang, et autres
Publié: (2026)
par: Jia, Haoxiang, et autres
Publié: (2026)
MANTRA: Enhancing Automated Method-Level Refactoring with Contextual RAG and Multi-Agent LLM Collaboration
par: Xu, Yisen, et autres
Publié: (2025)
par: Xu, Yisen, et autres
Publié: (2025)
Towards Structured, State-Aware, and Execution-Grounded Reasoning for Software Engineering Agents
par: Tse-Hsun, et autres
Publié: (2026)
par: Tse-Hsun, et autres
Publié: (2026)
BitsAI-CR: Automated Code Review via LLM in Practice
par: Sun, Tao, et autres
Publié: (2025)
par: Sun, Tao, et autres
Publié: (2025)
Static Program Slicing Using Language Models With Dataflow-Aware Pretraining and Constrained Decoding
par: He, Pengfei, et autres
Publié: (2026)
par: He, Pengfei, et autres
Publié: (2026)
Empowering AIOps: Leveraging Large Language Models for IT Operations Management
par: Vitui, Arthur, et autres
Publié: (2025)
par: Vitui, Arthur, et autres
Publié: (2025)
On Rank Aggregating Test Prioritizations
par: Mondal, Shouvick, et autres
Publié: (2024)
par: Mondal, Shouvick, et autres
Publié: (2024)
CI-Repair-Bench: A Repository-Aware Benchmark for Automated Patch Validation via CI Workflows
par: Muna, Rabeya Khatun, et autres
Publié: (2026)
par: Muna, Rabeya Khatun, et autres
Publié: (2026)
LLMParser: An Exploratory Study on Using Large Language Models for Log Parsing
par: Ma, Zeyang, et autres
Publié: (2024)
par: Ma, Zeyang, et autres
Publié: (2024)
Evaluating LLM-Generated Code: A Benchmark and Developer Study
par: Szych, Joanna, et autres
Publié: (2026)
par: Szych, Joanna, et autres
Publié: (2026)
A First Look at the Self-Admitted Technical Debt in Test Code: Taxonomy and Detection
par: Islam, Shahidul, et autres
Publié: (2025)
par: Islam, Shahidul, et autres
Publié: (2025)
Are Benchmark Tests Strong Enough? Mutation-Guided Diagnosis and Augmentation of Regression Suites
par: Li, Chenglin, et autres
Publié: (2026)
par: Li, Chenglin, et autres
Publié: (2026)
Automated Repair of Ambiguous Problem Descriptions for LLM-Based Code Generation
par: Jia, Haoxiang, et autres
Publié: (2025)
par: Jia, Haoxiang, et autres
Publié: (2025)
Code Refactoring with LLM: A Comprehensive Evaluation With Few-Shot Settings
par: Tapader, Md. Raihan, et autres
Publié: (2025)
par: Tapader, Md. Raihan, et autres
Publié: (2025)
Order Matters! An Empirical Study on Large Language Models' Input Order Bias in Software Fault Localization
par: Rafi, Md Nakhla, et autres
Publié: (2024)
par: Rafi, Md Nakhla, et autres
Publié: (2024)
Modern Code Reviews -- Survey of Literature and Practice
par: Badampudi, Deepika, et autres
Publié: (2024)
par: Badampudi, Deepika, et autres
Publié: (2024)
RobuNFR: Evaluating the Robustness of Large Language Models on Non-Functional Requirements Aware Code Generation
par: Lin, Feng, et autres
Publié: (2025)
par: Lin, Feng, et autres
Publié: (2025)
Studying and Benchmarking Large Language Models For Log Level Suggestion
par: Heng, Yi Wen, et autres
Publié: (2024)
par: Heng, Yi Wen, et autres
Publié: (2024)
Back to the Future! Studying Data Cleanness in Defects4J and its Impact on Fault Localization
par: Rafi, Md Nakhla, et autres
Publié: (2023)
par: Rafi, Md Nakhla, et autres
Publié: (2023)
Copilot Arena: A Platform for Code LLM Evaluation in the Wild
par: Chi, Wayne, et autres
Publié: (2025)
par: Chi, Wayne, et autres
Publié: (2025)
Practical Program Repair in the Era of Large Pre-trained Language Models
par: Xia, Chunqiu Steven, et autres
Publié: (2022)
par: Xia, Chunqiu Steven, et autres
Publié: (2022)
An Empirical Study on the Characteristics of Database Access Bugs in Java Applications
par: Liu, Wei, et autres
Publié: (2024)
par: Liu, Wei, et autres
Publié: (2024)
Evaluating and Achieving Controllable Code Completion in Code LLM
par: Zhang, Jiajun, et autres
Publié: (2026)
par: Zhang, Jiajun, et autres
Publié: (2026)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
par: Pan, Zhiyuan, et autres
Publié: (2025)
par: Pan, Zhiyuan, et autres
Publié: (2025)
Documents similaires
-
CODEPROMPTZIP: Code-specific Prompt Compression for Retrieval-Augmented Generation in Coding Tasks with LMs
par: He, Pengfei, et autres
Publié: (2025) -
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
par: Xu, Yisen, et autres
Publié: (2026) -
Evaluating the Effectiveness and Efficiency of Demonstration Retrievers in RAG for Coding Tasks
par: He, Pengfei, et autres
Publié: (2024) -
SLICET5: Static Program Slicing using Language Models with Copy Mechanism and Constrained Decoding
par: He, Pengfei, et autres
Publié: (2025) -
Identifying Performance-Sensitive Configurations in Software Systems through Code Analysis with LLM Agents
par: Wang, Zehao, et autres
Publié: (2024)