Evolution without an Oracle: Driving Effective Evolution with LLM Judges
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Zhe, Yang, Yuheng, Wen, Haibin, Qiu, Xiaojie, Zhang, Zaixi, Zhang, Qingfu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Understanding to Excelling: Template-Free Algorithm Design through Structural-Functional Co-Evolution
by: Zhao, Zhe, et al.
Published: (2025)
by: Zhao, Zhe, et al.
Published: (2025)
Evolution of Kernels: Automated RISC-V Kernel Optimization with Large Language Models
by: Chen, Siyuan, et al.
Published: (2025)
by: Chen, Siyuan, et al.
Published: (2025)
LLM-Driven Kernel Evolution: Automating Driver Updates in Linux
by: Kharlamova, Arina, et al.
Published: (2025)
by: Kharlamova, Arina, et al.
Published: (2025)
Effective LLM Code Refinement via Property-Oriented and Structurally Minimal Feedback
by: He, Lehan, et al.
Published: (2025)
by: He, Lehan, et al.
Published: (2025)
Go-Oracle: Automated Test Oracle for Go Concurrency Bugs
by: Tsimpourlas, Foivos, et al.
Published: (2024)
by: Tsimpourlas, Foivos, et al.
Published: (2024)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
by: Zhang, Wentao, et al.
Published: (2026)
by: Zhang, Wentao, et al.
Published: (2026)
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering
by: Zhao, Zixiao, et al.
Published: (2026)
by: Zhao, Zixiao, et al.
Published: (2026)
Evaluating LLM-Based Test Generation Under Software Evolution
by: Haroon, Sabaat, et al.
Published: (2026)
by: Haroon, Sabaat, et al.
Published: (2026)
irace-evo: Automatic Algorithm Configuration Extended With LLM-Based Code Evolution
by: Sartori, Camilo Chacón, et al.
Published: (2025)
by: Sartori, Camilo Chacón, et al.
Published: (2025)
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
AutoICE: Automatically Synthesizing Verifiable C Code via LLM-driven Evolution
by: Luo, Weilin, et al.
Published: (2025)
by: Luo, Weilin, et al.
Published: (2025)
EvoClaw: Evaluating AI Agents on Continuous Software Evolution
by: Deng, Gangda, et al.
Published: (2026)
by: Deng, Gangda, et al.
Published: (2026)
RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation
by: Liang, Linxi, et al.
Published: (2025)
by: Liang, Linxi, et al.
Published: (2025)
Engineering AI Judge Systems
by: Lin, Jiahuei, et al.
Published: (2024)
by: Lin, Jiahuei, et al.
Published: (2024)
Understanding LLM-Driven Test Oracle Generation
by: Bodicoat, Adam, et al.
Published: (2026)
by: Bodicoat, Adam, et al.
Published: (2026)
Automated Proof Generation for Rust Code via Self-Evolution
by: Chen, Tianyu, et al.
Published: (2024)
by: Chen, Tianyu, et al.
Published: (2024)
CODESYNC: Synchronizing Large Language Models with Dynamic Code Evolution at Scale
by: Wang, Chenlong, et al.
Published: (2025)
by: Wang, Chenlong, et al.
Published: (2025)
Loosely-Structured Software: Engineering Context, Structure, and Evolution Entropy in Runtime-Rewired Multi-Agent Systems
by: Zhang, Weihao, et al.
Published: (2026)
by: Zhang, Weihao, et al.
Published: (2026)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
Automated Snippet-Alignment Data Augmentation for Code Translation
by: Zhang, Zhiming, et al.
Published: (2025)
by: Zhang, Zhiming, et al.
Published: (2025)
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
by: Huang, Donghao, et al.
Published: (2025)
by: Huang, Donghao, et al.
Published: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
by: Jiang, Hongchao, et al.
Published: (2025)
by: Jiang, Hongchao, et al.
Published: (2025)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools
by: Son, Ha Min, et al.
Published: (2025)
by: Son, Ha Min, et al.
Published: (2025)
AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
LibEvolutionEval: A Benchmark and Study for Version-Specific Code Generation
by: Kuhar, Sachit, et al.
Published: (2024)
by: Kuhar, Sachit, et al.
Published: (2024)
Evolution of IVR building techniques: from code writing to AI-powered automation
by: Shaikh, Khushbu Mehboob, et al.
Published: (2024)
by: Shaikh, Khushbu Mehboob, et al.
Published: (2024)
From Queries to Insights: Agentic LLM Pipelines for Spatio-Temporal Text-to-SQL
by: Redd, Manu, et al.
Published: (2025)
by: Redd, Manu, et al.
Published: (2025)
On Simulation-Guided LLM-based Code Generation for Safe Autonomous Driving Software
by: Nouri, Ali, et al.
Published: (2025)
by: Nouri, Ali, et al.
Published: (2025)
CoCoEvo: Co-Evolution of Programs and Test Cases to Enhance Code Generation
by: Li, Kefan, et al.
Published: (2025)
by: Li, Kefan, et al.
Published: (2025)
Beyond Isolated Tasks: A Framework for Evaluating Coding Agents on Sequential Software Evolution
by: Shastry, KN Ajay, et al.
Published: (2026)
by: Shastry, KN Ajay, et al.
Published: (2026)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
by: Lai, Peng, et al.
Published: (2026)
by: Lai, Peng, et al.
Published: (2026)
LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile Apps
by: Zhao, Shanhui, et al.
Published: (2025)
by: Zhao, Shanhui, et al.
Published: (2025)
AutoDroid: LLM-powered Task Automation in Android
by: Wen, Hao, et al.
Published: (2023)
by: Wen, Hao, et al.
Published: (2023)
OrcaLoca: An LLM Agent Framework for Software Issue Localization
by: Yu, Zhongming, et al.
Published: (2025)
by: Yu, Zhongming, et al.
Published: (2025)
ReCode: Improving LLM-based Code Repair with Fine-Grained Retrieval-Augmented Generation
by: Zhao, Yicong, et al.
Published: (2025)
by: Zhao, Yicong, et al.
Published: (2025)
From Translation to Superset: Benchmark-Driven Evolution of a Production AI Agent from Rust to Python
by: Wang, Jinhua, et al.
Published: (2026)
by: Wang, Jinhua, et al.
Published: (2026)
Vul-R2: A Reasoning LLM for Automated Vulnerability Repair
by: Wen, Xin-Cheng, et al.
Published: (2025)
by: Wen, Xin-Cheng, et al.
Published: (2025)
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
by: Ma, Ming, et al.
Published: (2025)
by: Ma, Ming, et al.
Published: (2025)
Similar Items
-
From Understanding to Excelling: Template-Free Algorithm Design through Structural-Functional Co-Evolution
by: Zhao, Zhe, et al.
Published: (2025) -
Evolution of Kernels: Automated RISC-V Kernel Optimization with Large Language Models
by: Chen, Siyuan, et al.
Published: (2025) -
LLM-Driven Kernel Evolution: Automating Driver Updates in Linux
by: Kharlamova, Arina, et al.
Published: (2025) -
Effective LLM Code Refinement via Property-Oriented and Structurally Minimal Feedback
by: He, Lehan, et al.
Published: (2025) -
Go-Oracle: Automated Test Oracle for Go Concurrency Bugs
by: Tsimpourlas, Foivos, et al.
Published: (2024)