LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Donghao, Chew, Shila, Dutkiewicz, Anna, Wang, Zhaoxia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
por: Weng, Shihao, et al.
Publicado: (2026)
por: Weng, Shihao, et al.
Publicado: (2026)
Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs
por: Yang, Ya-Ting, et al.
Publicado: (2026)
por: Yang, Ya-Ting, et al.
Publicado: (2026)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
por: Wang, Ruiqi, et al.
Publicado: (2025)
por: Wang, Ruiqi, et al.
Publicado: (2025)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
por: Gao, Shuzheng, et al.
Publicado: (2025)
por: Gao, Shuzheng, et al.
Publicado: (2025)
PARCER as an Operational Contract to Reduce Variance, Cost, and Risk in LLM Systems
por: Filho, Elzo Brito dos Santos
Publicado: (2026)
por: Filho, Elzo Brito dos Santos
Publicado: (2026)
Towards Reliable LLM-Driven Fuzz Testing: Vision and Road Ahead
por: Cheng, Yiran, et al.
Publicado: (2025)
por: Cheng, Yiran, et al.
Publicado: (2025)
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
por: Le-Cong, Thanh, et al.
Publicado: (2024)
por: Le-Cong, Thanh, et al.
Publicado: (2024)
MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems
por: Jia, Jin, et al.
Publicado: (2026)
por: Jia, Jin, et al.
Publicado: (2026)
FuzzAug: Data Augmentation by Coverage-guided Fuzzing for Neural Test Generation
por: He, Yifeng, et al.
Publicado: (2024)
por: He, Yifeng, et al.
Publicado: (2024)
Synergizing Code Coverage and Gameplay Intent: Coverage-Aware Game Playtesting with LLM-Guided Reinforcement Learning
por: Mu, Enhong, et al.
Publicado: (2025)
por: Mu, Enhong, et al.
Publicado: (2025)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
por: Lai, Peng, et al.
Publicado: (2026)
por: Lai, Peng, et al.
Publicado: (2026)
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering
por: Zhao, Zixiao, et al.
Publicado: (2026)
por: Zhao, Zixiao, et al.
Publicado: (2026)
AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
por: Wang, Ruiqi, et al.
Publicado: (2025)
por: Wang, Ruiqi, et al.
Publicado: (2025)
Evolution without an Oracle: Driving Effective Evolution with LLM Judges
por: Zhao, Zhe, et al.
Publicado: (2025)
por: Zhao, Zhe, et al.
Publicado: (2025)
WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
por: Li, Chunyang, et al.
Publicado: (2025)
por: Li, Chunyang, et al.
Publicado: (2025)
Evaluating LLM-Based Test Generation Under Software Evolution
por: Haroon, Sabaat, et al.
Publicado: (2026)
por: Haroon, Sabaat, et al.
Publicado: (2026)
Klear-CodeTest: Scalable Test Case Generation for Code Reinforcement Learning
por: Fu, Jia, et al.
Publicado: (2025)
por: Fu, Jia, et al.
Publicado: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
por: Jiang, Hongchao, et al.
Publicado: (2025)
por: Jiang, Hongchao, et al.
Publicado: (2025)
Beyond Accuracy: An Empirical Study on Unit Testing in Open-source Deep Learning Projects
por: Wang, Han, et al.
Publicado: (2024)
por: Wang, Han, et al.
Publicado: (2024)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
por: Zhou, Xin, et al.
Publicado: (2025)
por: Zhou, Xin, et al.
Publicado: (2025)
Evaluation-Driven Development and Operations of LLM Agents: A Process Model and Reference Architecture
por: Xia, Boming, et al.
Publicado: (2024)
por: Xia, Boming, et al.
Publicado: (2024)
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
por: Oliva, Gustavo A., et al.
Publicado: (2025)
por: Oliva, Gustavo A., et al.
Publicado: (2025)
AgentGit: A Version Control Framework for Reliable and Scalable LLM-Powered Multi-Agent Systems
por: Li, Yang, et al.
Publicado: (2025)
por: Li, Yang, et al.
Publicado: (2025)
Reducing Cost of LLM Agents with Trajectory Reduction
por: Xiao, Yuan-An, et al.
Publicado: (2025)
por: Xiao, Yuan-An, et al.
Publicado: (2025)
Bias Testing and Mitigation in LLM-based Code Generation
por: Huang, Dong, et al.
Publicado: (2023)
por: Huang, Dong, et al.
Publicado: (2023)
VerilogReader: LLM-Aided Hardware Test Generation
por: Ma, Ruiyang, et al.
Publicado: (2024)
por: Ma, Ruiyang, et al.
Publicado: (2024)
Beyond Accuracy: LLM Variability in Evidence Screening for Software Engineering SLRs
por: Hida, Gilberto Sussumu, et al.
Publicado: (2026)
por: Hida, Gilberto Sussumu, et al.
Publicado: (2026)
Validating LLM-Generated Programs with Metamorphic Prompt Testing
por: Wang, Xiaoyin, et al.
Publicado: (2024)
por: Wang, Xiaoyin, et al.
Publicado: (2024)
InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution
por: Li, KeFan, et al.
Publicado: (2025)
por: Li, KeFan, et al.
Publicado: (2025)
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
por: Zhou, Zenghui, et al.
Publicado: (2026)
por: Zhou, Zenghui, et al.
Publicado: (2026)
A Self-Healing Framework for Reliable LLM-Based Autonomous Agents
por: Jeong, Cheonsu, et al.
Publicado: (2026)
por: Jeong, Cheonsu, et al.
Publicado: (2026)
ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization
por: Huang, Yixu, et al.
Publicado: (2026)
por: Huang, Yixu, et al.
Publicado: (2026)
Fuzzy Inference System for Test Case Prioritization in Software Testing
por: Karatayev, Aron, et al.
Publicado: (2024)
por: Karatayev, Aron, et al.
Publicado: (2024)
Engineering AI Judge Systems
por: Lin, Jiahuei, et al.
Publicado: (2024)
por: Lin, Jiahuei, et al.
Publicado: (2024)
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
por: Cui, Yi
Publicado: (2025)
por: Cui, Yi
Publicado: (2025)
Call-Chain-Aware LLM-Based Test Generation for Java Projects
por: Wang, Guancheng, et al.
Publicado: (2026)
por: Wang, Guancheng, et al.
Publicado: (2026)
Methodological Framework for Quantifying Semantic Test Coverage in RAG Systems
por: Broestl, Noah, et al.
Publicado: (2025)
por: Broestl, Noah, et al.
Publicado: (2025)
RESTestBench: A Benchmark for Evaluating the Effectiveness of LLM-Generated REST API Test Cases from NL Requirements
por: Kogler, Leon, et al.
Publicado: (2026)
por: Kogler, Leon, et al.
Publicado: (2026)
Graph-Based Self-Healing Tool Routing for Cost-Efficient LLM Agents
por: Bholani, Neeraj
Publicado: (2026)
por: Bholani, Neeraj
Publicado: (2026)
Can LLM Generate Regression Tests for Software Commits?
por: Liu, Jing, et al.
Publicado: (2025)
por: Liu, Jing, et al.
Publicado: (2025)
Ejemplares similares
-
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
por: Weng, Shihao, et al.
Publicado: (2026) -
Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs
por: Yang, Ya-Ting, et al.
Publicado: (2026) -
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
por: Wang, Ruiqi, et al.
Publicado: (2025) -
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
por: Gao, Shuzheng, et al.
Publicado: (2025) -
PARCER as an Operational Contract to Reduce Variance, Cost, and Risk in LLM Systems
por: Filho, Elzo Brito dos Santos
Publicado: (2026)