Understanding the Limits of Automated Evaluation for Code Review Bots in Practice
Fuente:
arXiv
Saved in:
| Main Authors: | Karakaya, Veli, Torun, Utku Boran, Uçar, Baykal Mehmet, Tüzün, Eray |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluation of LLM-Based Software Engineering Tools: Practices, Challenges, and Future Directions
by: Torun, Utku Boran, et al.
Published: (2026)
by: Torun, Utku Boran, et al.
Published: (2026)
Past, Present, and Future of Bug Tracking in the Generative AI Era
by: Torun, Utku Boran, et al.
Published: (2025)
by: Torun, Utku Boran, et al.
Published: (2025)
Towards Automated Detection of Inline Code Comment Smells
by: Oztas, Ipek, et al.
Published: (2025)
by: Oztas, Ipek, et al.
Published: (2025)
Automated Code Review In Practice
by: Cihan, Umut, et al.
Published: (2024)
by: Cihan, Umut, et al.
Published: (2024)
Evaluating Large Language Models for Code Review
by: Cihan, Umut, et al.
Published: (2025)
by: Cihan, Umut, et al.
Published: (2025)
Rethinking Code Review in the Age of AI: A Vision for Agentic Code Review
by: Kamalı, Hüseyin Özgür, et al.
Published: (2026)
by: Kamalı, Hüseyin Özgür, et al.
Published: (2026)
ImproBR: Bug Report Improver Using LLMs
by: Akyol, Emre Furkan, et al.
Published: (2026)
by: Akyol, Emre Furkan, et al.
Published: (2026)
Factors Influencing the Quality of AI-Generated Code: A Synthesis of Empirical Evidence
by: Geruslu, Vehid, et al.
Published: (2026)
by: Geruslu, Vehid, et al.
Published: (2026)
Agents in the Sandbox: End-to-End Crash Bug Reproduction for Minecraft
by: Yapağcı, Eray, et al.
Published: (2025)
by: Yapağcı, Eray, et al.
Published: (2025)
Automated Root-Cause Subclassification and No-Code Fix Generation for Invalid Bug Reports
by: Gon, Mahmut Furkan, et al.
Published: (2026)
by: Gon, Mahmut Furkan, et al.
Published: (2026)
Automated Classification of Human Code Review Comments with Large Language Models
by: Çağlar, Semih, et al.
Published: (2026)
by: Çağlar, Semih, et al.
Published: (2026)
LACY: Simulating Expert Mentoring for Software Onboarding with Code Tours
by: Kara, Zeynep Begüm, et al.
Published: (2026)
by: Kara, Zeynep Begüm, et al.
Published: (2026)
SmartDelta Methodology: Automated Quality Assurance and Optimization for Incremental System Engineering
by: Dornauer, Benedikt, et al.
Published: (2025)
by: Dornauer, Benedikt, et al.
Published: (2025)
Previously on... Automating Code Review
by: Heumüller, Robert, et al.
Published: (2025)
by: Heumüller, Robert, et al.
Published: (2025)
Evaluating the Impact of Data Cleaning on the Quality of Generated Pull Request Descriptions
by: Tire, Kutay, et al.
Published: (2025)
by: Tire, Kutay, et al.
Published: (2025)
DroidBot-GPT: GPT-powered UI Automation for Android
by: Wen, Hao, et al.
Published: (2023)
by: Wen, Hao, et al.
Published: (2023)
PR-Aware Automated Unit Test Generation: Challenges and Opportunities
by: Haratian, Vahid, et al.
Published: (2026)
by: Haratian, Vahid, et al.
Published: (2026)
AI-Assisted Assessment of Coding Practices in Modern Code Review
by: Vijayvergiya, Manushree, et al.
Published: (2024)
by: Vijayvergiya, Manushree, et al.
Published: (2024)
Chatbot-Based Assessment of Code Understanding in Automated Programming Assessment Systems
by: Frankford, Eduard, et al.
Published: (2026)
by: Frankford, Eduard, et al.
Published: (2026)
CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
by: Naik, Atharva, et al.
Published: (2024)
by: Naik, Atharva, et al.
Published: (2024)
Automated Code Review Using Large Language Models with Symbolic Reasoning
by: Icoz, Busra, et al.
Published: (2025)
by: Icoz, Busra, et al.
Published: (2025)
Benchmarking LLMs for Fine-Grained Code Review with Enriched Context in Practice
by: Hu, Ruida, et al.
Published: (2025)
by: Hu, Ruida, et al.
Published: (2025)
Towards Practical Defect-Focused Automated Code Review
by: Lu, Junyi, et al.
Published: (2025)
by: Lu, Junyi, et al.
Published: (2025)
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs
by: Li, Ziyu, et al.
Published: (2024)
by: Li, Ziyu, et al.
Published: (2024)
A Serious Game Approach to Introduce the Code Review Practice
by: Baris Ardic, et al.
Published: (2024)
by: Baris Ardic, et al.
Published: (2024)
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations
by: Tóth, Rebeka, et al.
Published: (2024)
by: Tóth, Rebeka, et al.
Published: (2024)
Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review
by: Paul, Debalina Ghosh, et al.
Published: (2024)
by: Paul, Debalina Ghosh, et al.
Published: (2024)
HalluJudge: A Reference-Free Hallucination Detection for Context Misalignment in Code Review Automation
by: Tantithamthavorn, Kla, et al.
Published: (2026)
by: Tantithamthavorn, Kla, et al.
Published: (2026)
Automating Patch Set Generation from Code Review Comments Using Large Language Models
by: Rahman, Tajmilur, et al.
Published: (2024)
by: Rahman, Tajmilur, et al.
Published: (2024)
FastCode: Fast and Cost-Efficient Code Understanding and Reasoning
by: Li, Zhonghang, et al.
Published: (2026)
by: Li, Zhonghang, et al.
Published: (2026)
CodeSSM: Towards State Space Models for Code Understanding
by: Verma, Shweta, et al.
Published: (2025)
by: Verma, Shweta, et al.
Published: (2025)
AACR-Bench: Evaluating Automatic Code Review with Holistic Repository-Level Context
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
MutaBot: A Mutation Testing Approach for Chatbots
by: Urrico, Michael Ferdinando, et al.
Published: (2024)
by: Urrico, Michael Ferdinando, et al.
Published: (2024)
Theory of Code Space: Do Code Agents Understand Software Architecture?
by: Sapunov, Grigory
Published: (2026)
by: Sapunov, Grigory
Published: (2026)
The Future of Generative AI in Software Engineering: A Vision from Industry and Academia in the European GENIUS Project
by: Gröpler, Robin, et al.
Published: (2025)
by: Gröpler, Robin, et al.
Published: (2025)
Augmenting Large Language Models with Static Code Analysis for Automated Code Quality Improvements
by: Abtahi, Seyed Moein, et al.
Published: (2025)
by: Abtahi, Seyed Moein, et al.
Published: (2025)
On The Impact of Merge Request Deviations on Code Review Practices
by: Kansab, Samah, et al.
Published: (2025)
by: Kansab, Samah, et al.
Published: (2025)
Automated Snippet-Alignment Data Augmentation for Code Translation
by: Zhang, Zhiming, et al.
Published: (2025)
by: Zhang, Zhiming, et al.
Published: (2025)
BugSpotter: Automated Generation of Code Debugging Exercises
by: Pădurean, Victor-Alexandru, et al.
Published: (2024)
by: Pădurean, Victor-Alexandru, et al.
Published: (2024)
Auto-SPT: Automating Semantic Preserving Transformations for Code
by: Hooda, Ashish, et al.
Published: (2025)
by: Hooda, Ashish, et al.
Published: (2025)
Similar Items
-
Evaluation of LLM-Based Software Engineering Tools: Practices, Challenges, and Future Directions
by: Torun, Utku Boran, et al.
Published: (2026) -
Past, Present, and Future of Bug Tracking in the Generative AI Era
by: Torun, Utku Boran, et al.
Published: (2025) -
Towards Automated Detection of Inline Code Comment Smells
by: Oztas, Ipek, et al.
Published: (2025) -
Automated Code Review In Practice
by: Cihan, Umut, et al.
Published: (2024) -
Evaluating Large Language Models for Code Review
by: Cihan, Umut, et al.
Published: (2025)