An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Xin, Kim, Kisub, Zhang, Ting, Weyssow, Martin, Gomes, Luis F., Yang, Guang, Liu, Kui, Xia, Xin, Lo, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of Code
by: Weyssow, Martin, et al.
Published: (2023)
by: Weyssow, Martin, et al.
Published: (2023)
Exploring Parameter-Efficient Fine-Tuning Techniques for Code Generation with Large Language Models
by: Weyssow, Martin, et al.
Published: (2023)
by: Weyssow, Martin, et al.
Published: (2023)
Multi-LLM Collaboration + Data-Centric Innovation = 2x Better Vulnerability Repair
by: Zhou, Xin, et al.
Published: (2024)
by: Zhou, Xin, et al.
Published: (2024)
CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences
by: Weyssow, Martin, et al.
Published: (2024)
by: Weyssow, Martin, et al.
Published: (2024)
A Functional Software Reference Architecture for LLM-Integrated Systems
by: Bucaioni, Alessio, et al.
Published: (2025)
by: Bucaioni, Alessio, et al.
Published: (2025)
iCodeReviewer: Improving Secure Code Review with Mixture of Prompts
by: Peng, Yun, et al.
Published: (2025)
by: Peng, Yun, et al.
Published: (2025)
Representation Learning for Stack Overflow Posts: How Far are We?
by: He, Junda, et al.
Published: (2023)
by: He, Junda, et al.
Published: (2023)
Curiosity-Driven Testing for Sequential Decision-Making Process
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks
by: Hu, Xing, et al.
Published: (2025)
by: Hu, Xing, et al.
Published: (2025)
Artificial Intelligence for Software Architecture: Literature Review and the Road Ahead
by: Bucaioni, Alessio, et al.
Published: (2025)
by: Bucaioni, Alessio, et al.
Published: (2025)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
PatchZero: Zero-Shot Automatic Patch Correctness Assessment
by: Zhou, Xin, et al.
Published: (2023)
by: Zhou, Xin, et al.
Published: (2023)
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation
by: Yang, Guang, et al.
Published: (2025)
by: Yang, Guang, et al.
Published: (2025)
A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
by: Liu, Shuhan, et al.
Published: (2026)
by: Liu, Shuhan, et al.
Published: (2026)
Large Language Model for Vulnerability Detection: Emerging Results and Future Directions
by: Zhou, Xin, et al.
Published: (2024)
by: Zhou, Xin, et al.
Published: (2024)
When Deep Learning Meets Information Retrieval-based Bug Localization: A Survey
by: Niu, Feifei, et al.
Published: (2025)
by: Niu, Feifei, et al.
Published: (2025)
Harnessing Large Language Models for Curated Code Reviews
by: Sghaier, Oussama Ben, et al.
Published: (2025)
by: Sghaier, Oussama Ben, et al.
Published: (2025)
Exploring the Capabilities of LLMs for Code Change Related Tasks
by: Fan, Lishui, et al.
Published: (2024)
by: Fan, Lishui, et al.
Published: (2024)
Bridging Expert Knowledge with Deep Learning Techniques for Just-In-Time Defect Prediction
by: Zhou, Xin, et al.
Published: (2024)
by: Zhou, Xin, et al.
Published: (2024)
Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models
by: Chang, Jianming, et al.
Published: (2025)
by: Chang, Jianming, et al.
Published: (2025)
LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Beyond Surface Similarity: Evaluating LLM-Based Test Refactorings with Structural and Semantic Awareness
by: Ouédraogo, Wendkûuni C., et al.
Published: (2025)
by: Ouédraogo, Wendkûuni C., et al.
Published: (2025)
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
by: Zhang, Chenchen, et al.
Published: (2025)
by: Zhang, Chenchen, et al.
Published: (2025)
Towards a Human-in-the-Loop Framework for Reliable Patch Evaluation Using an LLM-as-a-Judge
by: Shi, Sherry, et al.
Published: (2025)
by: Shi, Sherry, et al.
Published: (2025)
Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools
by: Son, Ha Min, et al.
Published: (2025)
by: Son, Ha Min, et al.
Published: (2025)
Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants
by: Richards, Jonan, et al.
Published: (2025)
by: Richards, Jonan, et al.
Published: (2025)
LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding
by: Amin, Md Faizul Ibne, et al.
Published: (2026)
by: Amin, Md Faizul Ibne, et al.
Published: (2026)
Rethinking Cognitive Complexity for Unit Tests: Toward a Readability-Aware Metric Grounded in Developer Perception
by: Ouédraogo, Wendkûuni C., et al.
Published: (2025)
by: Ouédraogo, Wendkûuni C., et al.
Published: (2025)
An Empirical Study of Vulnerable Package Dependencies in LLM Repositories
by: Liu, Shuhan, et al.
Published: (2025)
by: Liu, Shuhan, et al.
Published: (2025)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
by: Moon, Jiwon, et al.
Published: (2025)
by: Moon, Jiwon, et al.
Published: (2025)
VisDocSketcher: Towards Scalable Visual Documentation with Agentic Systems
by: Gomes, Luís F., et al.
Published: (2025)
by: Gomes, Luís F., et al.
Published: (2025)
SolidCoder: Bridging the Mental-Reality Gap in LLM Code Generation through Concrete Execution
by: Lee, Woojin, et al.
Published: (2026)
by: Lee, Woojin, et al.
Published: (2026)
4D-ARE: Bridging the Attribution Gap in LLM Agent Requirements Engineering
by: Yu, Bo, et al.
Published: (2026)
by: Yu, Bo, et al.
Published: (2026)
Towards Bridging Language Gaps in OSS with LLM-Driven Documentation Translation
by: Adejumo, Elijah Kayode, et al.
Published: (2025)
by: Adejumo, Elijah Kayode, et al.
Published: (2025)
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
Integrating Rules and Semantics for LLM-Based C-to-Rust Translation
by: Luo, Feng, et al.
Published: (2025)
by: Luo, Feng, et al.
Published: (2025)
AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
An Empirical Study of Speculative Decoding on Software Engineering Tasks
by: Li, Yijia, et al.
Published: (2026)
by: Li, Yijia, et al.
Published: (2026)
Benchmarking Large Language Models for Multi-Language Software Vulnerability Detection
by: Zhang, Ting, et al.
Published: (2025)
by: Zhang, Ting, et al.
Published: (2025)
SLICEMATE: Accurate and Scalable Static Program Slicing via LLM-Powered Agents
by: Chang, Jianming, et al.
Published: (2025)
by: Chang, Jianming, et al.
Published: (2025)
Similar Items
-
On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of Code
by: Weyssow, Martin, et al.
Published: (2023) -
Exploring Parameter-Efficient Fine-Tuning Techniques for Code Generation with Large Language Models
by: Weyssow, Martin, et al.
Published: (2023) -
Multi-LLM Collaboration + Data-Centric Innovation = 2x Better Vulnerability Repair
by: Zhou, Xin, et al.
Published: (2024) -
CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences
by: Weyssow, Martin, et al.
Published: (2024) -
A Functional Software Reference Architecture for LLM-Integrated Systems
by: Bucaioni, Alessio, et al.
Published: (2025)