Towards a Human-in-the-Loop Framework for Reliable Patch Evaluation Using an LLM-as-a-Judge
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Sherry, Wei, Renyao, Tufano, Michele, Cambronero, José, Cheng, Runxiang, Ivančić, Franjo, Rondon, Pat |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair
by: Cheng, Runxiang, et al.
Published: (2026)
by: Cheng, Runxiang, et al.
Published: (2026)
Agentic Bug Reproduction for Effective Automated Program Repair at Google
by: Cheng, Runxiang, et al.
Published: (2025)
by: Cheng, Runxiang, et al.
Published: (2025)
Abstain and Validate: A Dual-LLM Policy for Reducing Noise in Agentic Program Repair
by: Cambronero, José, et al.
Published: (2025)
by: Cambronero, José, et al.
Published: (2025)
Evaluating Agent-based Program Repair at Google
by: Rondon, Pat, et al.
Published: (2025)
by: Rondon, Pat, et al.
Published: (2025)
Towards Verified Code Reasoning by LLMs
by: Sistla, Meghana, et al.
Published: (2025)
by: Sistla, Meghana, et al.
Published: (2025)
LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding
by: Amin, Md Faizul Ibne, et al.
Published: (2026)
by: Amin, Md Faizul Ibne, et al.
Published: (2026)
KGym: A Platform and Dataset to Benchmark Large Language Models on Linux Kernel Crash Resolution
by: Mathai, Alex, et al.
Published: (2024)
by: Mathai, Alex, et al.
Published: (2024)
Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All
by: Huang, Chenxi, et al.
Published: (2026)
by: Huang, Chenxi, et al.
Published: (2026)
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
by: Huang, Donghao, et al.
Published: (2025)
by: Huang, Donghao, et al.
Published: (2025)
Towards Summarizing Code Snippets Using Pre-Trained Transformers
by: Mastropaolo, Antonio, et al.
Published: (2024)
by: Mastropaolo, Antonio, et al.
Published: (2024)
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering
by: Zhao, Zixiao, et al.
Published: (2026)
by: Zhao, Zixiao, et al.
Published: (2026)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
by: Agarwal, Anisha, et al.
Published: (2024)
by: Agarwal, Anisha, et al.
Published: (2024)
On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization
by: Crupi, Giuseppe, et al.
Published: (2025)
by: Crupi, Giuseppe, et al.
Published: (2025)
Automating Code Review: A Systematic Literature Review
by: Tufano, Rosalia, et al.
Published: (2025)
by: Tufano, Rosalia, et al.
Published: (2025)
Improving Code Generation via Small Language Model-as-a-judge
by: Crupi, Giuseppe, et al.
Published: (2026)
by: Crupi, Giuseppe, et al.
Published: (2026)
From Correctness to Collaboration: Toward a Human-Centered Framework for Evaluating AI Agent Behavior in Software Engineering
by: Dong, Tao, et al.
Published: (2025)
by: Dong, Tao, et al.
Published: (2025)
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
Reinforcement Learning from Automatic Feedback for High-Quality Unit Test Generation
by: Steenhoek, Benjamin, et al.
Published: (2023)
by: Steenhoek, Benjamin, et al.
Published: (2023)
Reinforcement Learning from Automatic Feedback for High-Quality Unit Test Generation
by: Steenhoek, Benjamin, et al.
Published: (2024)
by: Steenhoek, Benjamin, et al.
Published: (2024)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
SEART Data Hub: Streamlining Large-Scale Source Code Mining and Pre-Processing
by: Dabić, Ozren, et al.
Published: (2024)
by: Dabić, Ozren, et al.
Published: (2024)
Studying Quality Improvements Recommended via Manual and Automated Code Review
by: Crupi, Giuseppe, et al.
Published: (2026)
by: Crupi, Giuseppe, et al.
Published: (2026)
Towards Reliable LLM-Driven Fuzz Testing: Vision and Road Ahead
by: Cheng, Yiran, et al.
Published: (2025)
by: Cheng, Yiran, et al.
Published: (2025)
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
Learning From Developers: Towards Reliable Patch Validation at Scale for Linux
by: Lin, Chih-En, et al.
Published: (2026)
by: Lin, Chih-En, et al.
Published: (2026)
Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation
by: Khati, Dipin, et al.
Published: (2025)
by: Khati, Dipin, et al.
Published: (2025)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
by: Lai, Peng, et al.
Published: (2026)
by: Lai, Peng, et al.
Published: (2026)
Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge
by: Mu, Wenhan, et al.
Published: (2025)
by: Mu, Wenhan, et al.
Published: (2025)
LLM4VV: Exploring LLM-as-a-Judge for Validation and Verification Testsuites
by: Sollenberger, Zachariah, et al.
Published: (2024)
by: Sollenberger, Zachariah, et al.
Published: (2024)
CrashFixer: A crash resolution agent for the Linux kernel
by: Mathai, Alex, et al.
Published: (2025)
by: Mathai, Alex, et al.
Published: (2025)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
by: Moon, Jiwon, et al.
Published: (2025)
by: Moon, Jiwon, et al.
Published: (2025)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
On the Generalizability of Transformer Models to Code Completions of Different Lengths
by: Cooper, Nathan, et al.
Published: (2025)
by: Cooper, Nathan, et al.
Published: (2025)
Towards Demystifying and Repairing LLM-in-the-Loop Vulnerabilities
by: Ma, Yujie, et al.
Published: (2026)
by: Ma, Yujie, et al.
Published: (2026)
Is LLM-Generated Code More Maintainable \& Reliable than Human-Written Code?
by: Molison, Alfred Santa, et al.
Published: (2025)
by: Molison, Alfred Santa, et al.
Published: (2025)
Tricky$^2$: Towards a Benchmark for Evaluating Human and LLM Error Interactions
by: Granger, Cole, et al.
Published: (2026)
by: Granger, Cole, et al.
Published: (2026)
AutoDev: Automated AI-Driven Development
by: Tufano, Michele, et al.
Published: (2024)
by: Tufano, Michele, et al.
Published: (2024)
Using a Feedback Loop for LLM-based Infrastructure as Code Generation
by: Palavalli, Mayur Amarnath, et al.
Published: (2024)
by: Palavalli, Mayur Amarnath, et al.
Published: (2024)
A Prompt-Based Framework for Loop Vulnerability Detection Using Local LLMs
by: Adeseye, Adeyemi, et al.
Published: (2026)
by: Adeseye, Adeyemi, et al.
Published: (2026)
PoC-Gym: Towards More Reliable LLM-Assisted Proof-of-Concept Exploit Generation
by: Gezgin, Derin, et al.
Published: (2026)
by: Gezgin, Derin, et al.
Published: (2026)
Similar Items
-
Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair
by: Cheng, Runxiang, et al.
Published: (2026) -
Agentic Bug Reproduction for Effective Automated Program Repair at Google
by: Cheng, Runxiang, et al.
Published: (2025) -
Abstain and Validate: A Dual-LLM Policy for Reducing Noise in Agentic Program Repair
by: Cambronero, José, et al.
Published: (2025) -
Evaluating Agent-based Program Repair at Google
by: Rondon, Pat, et al.
Published: (2025) -
Towards Verified Code Reasoning by LLMs
by: Sistla, Meghana, et al.
Published: (2025)