Automated Validation of LLM-based Evaluators for Software Engineering Artifacts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fandina, Ora Nova, Farchi, Eitan, Froimovich, Shmulik, Katan, Rami, Podolsky, Alice, Raz, Orna, Ziv, Avi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909721870467072
author Fandina, Ora Nova
Farchi, Eitan
Froimovich, Shmulik
Katan, Rami
Podolsky, Alice
Raz, Orna
Ziv, Avi
author_facet Fandina, Ora Nova
Farchi, Eitan
Froimovich, Shmulik
Katan, Rami
Podolsky, Alice
Raz, Orna
Ziv, Avi
contents Automation in software engineering increasingly relies on large language models (LLMs) to generate, review, and assess code artifacts. However, establishing LLMs as reliable evaluators remains an open challenge: human evaluations are costly, subjective and non scalable, while existing automated methods fail to discern fine grained variations in artifact quality. We introduce REFINE (Ranking Evaluators for FIne grained Nuanced Evaluation), an automated framework for benchmarking LLM based evaluators across software engineering tasks. REFINE comprises of two modules: Hierarchy Dataset Builder applies novel generation techniques to automatically synthesize artifacts with progressively reduced quality, and Evaluator Tester quantifies each candidate evaluator configuration by measuring how closely its rankings align with expected ordering. A key feature of REFINE is controllability: users can tune the granularity of degradation to progressively refine evaluator configurations, from coarse filtering to stress testing on subtle quality gaps. While the methodology is general, we focus on coding tasks reflecting the practical demands in our production setting. REFINE was integrated into IBM's internal development workflows and applied to code generation, translation, and summarization for COBOL, an enterprise critical programming language, using industrial data. It was used to identify LLM as a Judge configurations that lifted alignment scores from below $0.7$ to above $0.9$ in some coding tasks. These nuance sensitive evaluators are now actively used by model training teams to support model release decisions.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02827
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
Fandina, Ora Nova
Farchi, Eitan
Froimovich, Shmulik
Katan, Rami
Podolsky, Alice
Raz, Orna
Ziv, Avi
Software Engineering
Artificial Intelligence
Automation in software engineering increasingly relies on large language models (LLMs) to generate, review, and assess code artifacts. However, establishing LLMs as reliable evaluators remains an open challenge: human evaluations are costly, subjective and non scalable, while existing automated methods fail to discern fine grained variations in artifact quality. We introduce REFINE (Ranking Evaluators for FIne grained Nuanced Evaluation), an automated framework for benchmarking LLM based evaluators across software engineering tasks. REFINE comprises of two modules: Hierarchy Dataset Builder applies novel generation techniques to automatically synthesize artifacts with progressively reduced quality, and Evaluator Tester quantifies each candidate evaluator configuration by measuring how closely its rankings align with expected ordering. A key feature of REFINE is controllability: users can tune the granularity of degradation to progressively refine evaluator configurations, from coarse filtering to stress testing on subtle quality gaps. While the methodology is general, we focus on coding tasks reflecting the practical demands in our production setting. REFINE was integrated into IBM's internal development workflows and applied to code generation, translation, and summarization for COBOL, an enterprise critical programming language, using industrial data. It was used to identify LLM as a Judge configurations that lifted alignment scores from below $0.7$ to above $0.9$ in some coding tasks. These nuance sensitive evaluators are now actively used by model training teams to support model release decisions.
title Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2508.02827