LLM4VV: Exploring LLM-as-a-Judge for Validation and Verification Testsuites
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sollenberger, Zachariah, Patel, Jay, Munley, Christian, Jarmusch, Aaron, Chandrasekaran, Sunita |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM4VV: Developing LLM-Driven Testsuite for Compiler Validation
von: Munley, Christian, et al.
Veröffentlicht: (2023)
von: Munley, Christian, et al.
Veröffentlicht: (2023)
LLM4VV: Evaluating Cutting-Edge LLMs for Generation and Evaluation of Directive-Based Parallel Programming Model Compiler Tests
von: Sollenberger, Zachariah, et al.
Veröffentlicht: (2025)
von: Sollenberger, Zachariah, et al.
Veröffentlicht: (2025)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
von: Moon, Jiwon, et al.
Veröffentlicht: (2025)
von: Moon, Jiwon, et al.
Veröffentlicht: (2025)
LLM-as-a-Judge for Reference-less Automatic Code Validation and Refinement for Natural Language to Bash in IT Automation
von: Vo, Ngoc Phuoc An, et al.
Veröffentlicht: (2025)
von: Vo, Ngoc Phuoc An, et al.
Veröffentlicht: (2025)
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
von: He, Junda, et al.
Veröffentlicht: (2025)
von: He, Junda, et al.
Veröffentlicht: (2025)
DafnyPro: LLM-Assisted Automated Verification for Dafny Programs
von: Banerjee, Debangshu, et al.
Veröffentlicht: (2026)
von: Banerjee, Debangshu, et al.
Veröffentlicht: (2026)
SACTOR: LLM-Driven Correct and Idiomatic C to Rust Translation with Static Analysis and FFI-Based Verification
von: Zhou, Tianyang, et al.
Veröffentlicht: (2025)
von: Zhou, Tianyang, et al.
Veröffentlicht: (2025)
Towards a Human-in-the-Loop Framework for Reliable Patch Evaluation Using an LLM-as-a-Judge
von: Shi, Sherry, et al.
Veröffentlicht: (2025)
von: Shi, Sherry, et al.
Veröffentlicht: (2025)
RE-oriented Model Development with LLM Support and Deduction-based Verification
von: Klimek, Radoslaw
Veröffentlicht: (2025)
von: Klimek, Radoslaw
Veröffentlicht: (2025)
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering
von: Zhao, Zixiao, et al.
Veröffentlicht: (2026)
von: Zhao, Zixiao, et al.
Veröffentlicht: (2026)
LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding
von: Amin, Md Faizul Ibne, et al.
Veröffentlicht: (2026)
von: Amin, Md Faizul Ibne, et al.
Veröffentlicht: (2026)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2026)
Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural Analysis
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2025)
von: Jarmusch, Aaron, et al.
Veröffentlicht: (2025)
CodeGrad: Integrating Multi-Step Verification with Gradient-Based LLM Refinement
von: Zhang, Yueke, et al.
Veröffentlicht: (2025)
von: Zhang, Yueke, et al.
Veröffentlicht: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
From Program Slices to Causal Clarity: Evaluating Faithful, Actionable LLM-Generated Failure Explanations via Context Partitioning and LLM-as-a-Judge
von: Porbeck, Julius, et al.
Veröffentlicht: (2026)
von: Porbeck, Julius, et al.
Veröffentlicht: (2026)
Challenges of Virtual Validation and Verification for Automotive Functions
von: Cabrero-Daniel, Beatriz, et al.
Veröffentlicht: (2025)
von: Cabrero-Daniel, Beatriz, et al.
Veröffentlicht: (2025)
Evolution without an Oracle: Driving Effective Evolution with LLM Judges
von: Zhao, Zhe, et al.
Veröffentlicht: (2025)
von: Zhao, Zhe, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
von: Weng, Shihao, et al.
Veröffentlicht: (2026)
von: Weng, Shihao, et al.
Veröffentlicht: (2026)
Agents4PLC: Automating Closed-loop PLC Code Generation and Verification in Industrial Control Systems using LLM-based Agents
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
LLM-Based Static Verification of Code Against Natural-Language Requirements: An Industrial Experience Report
von: Zhou, Zhi Quan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhi Quan, et al.
Veröffentlicht: (2026)
SpecSyn: LLM-based Synthesis and Refinement of Formal Specifications for Real-world Program Verification
von: Ma, Lezhi, et al.
Veröffentlicht: (2026)
von: Ma, Lezhi, et al.
Veröffentlicht: (2026)
LLM-Assisted Tool for Joint Generation of Formulas and Functions in Rule-Based Verification of Map Transformations
von: He, Ruidi, et al.
Veröffentlicht: (2025)
von: He, Ruidi, et al.
Veröffentlicht: (2025)
Context Conquers Parameters: Outperforming Proprietary LLM in Commit Message Generation
von: Imani, Aaron, et al.
Veröffentlicht: (2024)
von: Imani, Aaron, et al.
Veröffentlicht: (2024)
Exploring Creativity in Human-Human-LLM Collaborative Software Design
von: Jackson, Victoria, et al.
Veröffentlicht: (2026)
von: Jackson, Victoria, et al.
Veröffentlicht: (2026)
Using Assurance Cases to Guide Verification and Validation of Research Software
von: Smith, W. Spencer, et al.
Veröffentlicht: (2024)
von: Smith, W. Spencer, et al.
Veröffentlicht: (2024)
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
von: Huang, Donghao, et al.
Veröffentlicht: (2025)
von: Huang, Donghao, et al.
Veröffentlicht: (2025)
Verification and Validation of Autonomous Systems
von: Shetiya, Sneha Sudhir, et al.
Veröffentlicht: (2024)
von: Shetiya, Sneha Sudhir, et al.
Veröffentlicht: (2024)
Verification Limits Code LLM Training
von: Gureja, Srishti, et al.
Veröffentlicht: (2025)
von: Gureja, Srishti, et al.
Veröffentlicht: (2025)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
LogSage: An LLM-Based Framework for CI/CD Failure Detection and Remediation with Industrial Validation
von: Xu, Weiyuan, et al.
Veröffentlicht: (2025)
von: Xu, Weiyuan, et al.
Veröffentlicht: (2025)
Solver-Aided Verification of Policy Compliance in Tool-Augmented LLM Agents
von: Winston, Cailin, et al.
Veröffentlicht: (2026)
von: Winston, Cailin, et al.
Veröffentlicht: (2026)
Verifying LLM-Generated Code in the Context of Software Verification with Ada/SPARK
von: Cramer, Marcos, et al.
Veröffentlicht: (2025)
von: Cramer, Marcos, et al.
Veröffentlicht: (2025)
Exploring and Lifting the Robustness of LLM-powered Automated Program Repair with Metamorphic Testing
von: Xue, Pengyu, et al.
Veröffentlicht: (2024)
von: Xue, Pengyu, et al.
Veröffentlicht: (2024)
Understanding the Characteristics of LLM-Generated Property-Based Tests in Exploring Edge Cases
von: Tanaka, Hidetake, et al.
Veröffentlicht: (2025)
von: Tanaka, Hidetake, et al.
Veröffentlicht: (2025)
LLM4FP: LLM-Based Program Generation for Triggering Floating-Point Inconsistencies Across Compilers
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
Validating LLM-Generated Programs with Metamorphic Prompt Testing
von: Wang, Xiaoyin, et al.
Veröffentlicht: (2024)
von: Wang, Xiaoyin, et al.
Veröffentlicht: (2024)
Testing, Evaluation, Verification and Validation (TEVV) of Digital Twins: A Comprehensive Framework
von: Waters, Gabriella
Veröffentlicht: (2025)
von: Waters, Gabriella
Veröffentlicht: (2025)
AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
Evaluate-and-Purify: Fortifying Code Language Models Against Adversarial Attacks Using LLM-as-a-Judge
von: Mu, Wenhan, et al.
Veröffentlicht: (2025)
von: Mu, Wenhan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLM4VV: Developing LLM-Driven Testsuite for Compiler Validation
von: Munley, Christian, et al.
Veröffentlicht: (2023) -
LLM4VV: Evaluating Cutting-Edge LLMs for Generation and Evaluation of Directive-Based Parallel Programming Model Compiler Tests
von: Sollenberger, Zachariah, et al.
Veröffentlicht: (2025) -
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
von: Moon, Jiwon, et al.
Veröffentlicht: (2025) -
LLM-as-a-Judge for Reference-less Automatic Code Validation and Refinement for Natural Language to Bash in IT Automation
von: Vo, Ngoc Phuoc An, et al.
Veröffentlicht: (2025) -
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
von: He, Junda, et al.
Veröffentlicht: (2025)