Understanding LLM-Driven Test Oracle Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Bodicoat, Adam, Jahangirova, Gunel, Terragni, Valerio |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMLOOP: Improving LLM-Generated Code and Tests through Automated Iterative Feedback Loops
by: Ravi, Ravin, et al.
Published: (2026)
by: Ravi, Ravin, et al.
Published: (2026)
Real Faults in Deep Learning Fault Benchmarks: How Real Are They?
by: Jahangirova, Gunel, et al.
Published: (2024)
by: Jahangirova, Gunel, et al.
Published: (2024)
An Empirical Study of Fault Localisation Techniques for Deep Learning
by: Humbatova, Nargiz, et al.
Published: (2024)
by: Humbatova, Nargiz, et al.
Published: (2024)
GenMorph: Automatically Generating Metamorphic Relations via Genetic Programming
by: Ayerdi, Jon, et al.
Published: (2023)
by: Ayerdi, Jon, et al.
Published: (2023)
Metamorphic Testing of Large Language Models for Natural Language Processing
by: Cho, Steven, et al.
Published: (2025)
by: Cho, Steven, et al.
Published: (2025)
From Untestable to Testable: Metamorphic Testing in the Age of LLMs
by: Terragni, Valerio
Published: (2026)
by: Terragni, Valerio
Published: (2026)
LLMORPH: Automated Metamorphic Testing of Large Language Models
by: Cho, Steven, et al.
Published: (2026)
by: Cho, Steven, et al.
Published: (2026)
Automated Trustworthiness Testing for Machine Learning Classifiers
by: Cho, Steven, et al.
Published: (2024)
by: Cho, Steven, et al.
Published: (2024)
Incoherence as Oracle-less Measure of Error in LLM-Based Code Generation
by: Valentin, Thomas, et al.
Published: (2025)
by: Valentin, Thomas, et al.
Published: (2025)
A Theoretical Analysis of Test-Driven Code Generation
by: Menet, Nicolas, et al.
Published: (2026)
by: Menet, Nicolas, et al.
Published: (2026)
MR-Coupler: Automated Metamorphic Test Generation via Functional Coupling Analysis
by: Xu, Congying, et al.
Published: (2026)
by: Xu, Congying, et al.
Published: (2026)
Mutation-Guided LLM-based Test Generation at Meta
by: Foster, Christopher, et al.
Published: (2025)
by: Foster, Christopher, et al.
Published: (2025)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
by: Xu, WeiZhe, et al.
Published: (2026)
by: Xu, WeiZhe, et al.
Published: (2026)
DeepKnowledge: Generalisation-Driven Deep Learning Testing
by: Missaoui, Sondess, et al.
Published: (2024)
by: Missaoui, Sondess, et al.
Published: (2024)
The Future of AI-Driven Software Engineering
by: Terragni, Valerio, et al.
Published: (2024)
by: Terragni, Valerio, et al.
Published: (2024)
Generative AI to Generate Test Data Generators
by: Baudry, Benoit, et al.
Published: (2024)
by: Baudry, Benoit, et al.
Published: (2024)
A Regression Framework for Understanding Prompt Component Impact on LLM Performance
by: Lauziere, Andrew, et al.
Published: (2026)
by: Lauziere, Andrew, et al.
Published: (2026)
MASTEST: A LLM-Based Multi-Agent System For RESTful API Tests
by: Han, Xiaoke, et al.
Published: (2025)
by: Han, Xiaoke, et al.
Published: (2025)
Code Generation by Differential Test Time Scaling
by: He, Yifeng, et al.
Published: (2026)
by: He, Yifeng, et al.
Published: (2026)
BitsAI-Fix: LLM-Driven Approach for Automated Lint Error Resolution in Practice
by: Li, Yuanpeng, et al.
Published: (2025)
by: Li, Yuanpeng, et al.
Published: (2025)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
by: Bodla, Krishna Vamshi, et al.
Published: (2025)
by: Bodla, Krishna Vamshi, et al.
Published: (2025)
Co-Located Tests, Better AI Code: How Test Syntax Structure Affects Foundation Model Code Generation
by: Jacopin, Éric
Published: (2026)
by: Jacopin, Éric
Published: (2026)
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
by: Bai, Yifan, et al.
Published: (2026)
by: Bai, Yifan, et al.
Published: (2026)
Using Quality Attribute Scenarios for ML Model Test Case Generation
by: Brower-Sinning, Rachel, et al.
Published: (2024)
by: Brower-Sinning, Rachel, et al.
Published: (2024)
Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
by: Jiang, Shan, et al.
Published: (2026)
by: Jiang, Shan, et al.
Published: (2026)
CLOVER: A Test Case Generation Benchmark with Coverage, Long-Context, and Verification
by: Xu, Jiacheng, et al.
Published: (2025)
by: Xu, Jiacheng, et al.
Published: (2025)
Protocol-Driven Development: Governing Generated Software Through Invariants and Continuous Evidence
by: He, Jun, et al.
Published: (2026)
by: He, Jun, et al.
Published: (2026)
Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
by: Manglik, Akshay, et al.
Published: (2026)
by: Manglik, Akshay, et al.
Published: (2026)
Go-Oracle: Automated Test Oracle for Go Concurrency Bugs
by: Tsimpourlas, Foivos, et al.
Published: (2024)
by: Tsimpourlas, Foivos, et al.
Published: (2024)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
by: Qiu, Ruizhong, et al.
Published: (2024)
by: Qiu, Ruizhong, et al.
Published: (2024)
Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study
by: Storhaug, André, et al.
Published: (2024)
by: Storhaug, André, et al.
Published: (2024)
ToolFactory: Automating Tool Generation by Leveraging LLM to Understand REST API Documentations
by: Ni, Xinyi, et al.
Published: (2025)
by: Ni, Xinyi, et al.
Published: (2025)
Improving LLM-Driven Test Generation by Learning from Mocking Information
by: Lee, Jamie, et al.
Published: (2026)
by: Lee, Jamie, et al.
Published: (2026)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
by: Andrews, Martin, et al.
Published: (2025)
by: Andrews, Martin, et al.
Published: (2025)
scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns
by: Samsonau, Sergey V.
Published: (2026)
by: Samsonau, Sergey V.
Published: (2026)
How Robust are LLM-Generated Library Imports? An Empirical Study using Stack Overflow
by: Latendresse, Jasmine, et al.
Published: (2025)
by: Latendresse, Jasmine, et al.
Published: (2025)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate
by: Liu, Hugh Xuechen, et al.
Published: (2026)
by: Liu, Hugh Xuechen, et al.
Published: (2026)
Can Search-Based Testing with Pareto Optimization Effectively Cover Failure-Revealing Test Inputs?
by: Sorokin, Lev, et al.
Published: (2024)
by: Sorokin, Lev, et al.
Published: (2024)
Automatic Detection of LLM-Generated Code: A Comparative Case Study of Contemporary Models Across Function and Class Granularities
by: Rahman, Musfiqur, et al.
Published: (2024)
by: Rahman, Musfiqur, et al.
Published: (2024)
Similar Items
-
LLMLOOP: Improving LLM-Generated Code and Tests through Automated Iterative Feedback Loops
by: Ravi, Ravin, et al.
Published: (2026) -
Real Faults in Deep Learning Fault Benchmarks: How Real Are They?
by: Jahangirova, Gunel, et al.
Published: (2024) -
An Empirical Study of Fault Localisation Techniques for Deep Learning
by: Humbatova, Nargiz, et al.
Published: (2024) -
GenMorph: Automatically Generating Metamorphic Relations via Genetic Programming
by: Ayerdi, Jon, et al.
Published: (2023) -
Metamorphic Testing of Large Language Models for Natural Language Processing
by: Cho, Steven, et al.
Published: (2025)