Do LLMs generate test oracles that capture the actual or the expected program behaviour?
Fuente:
arXiv
Saved in:
| Main Authors: | Konstantinou, Michael, Degiovanni, Renzo, Papadakis, Mike |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How well LLM-based test generation techniques perform with newer LLM versions?
by: Konstantinou, Michael, et al.
Published: (2026)
by: Konstantinou, Michael, et al.
Published: (2026)
YATE: The Role of Test Repair in LLM-Based Unit Test Generation
by: Konstantinou, Michael, et al.
Published: (2025)
by: Konstantinou, Michael, et al.
Published: (2025)
Bounded Synthesis of Synchronized Distributed Models from Lightweight Specifications
by: Castro, Pablo F., et al.
Published: (2025)
by: Castro, Pablo F., et al.
Published: (2025)
Evaluating LLMs for One-Shot Patching of Real and Artificial Vulnerabilities
by: Garg, Aayush, et al.
Published: (2025)
by: Garg, Aayush, et al.
Published: (2025)
Inferring Code Correctness from Specification
by: Florian, Tambon, et al.
Published: (2026)
by: Florian, Tambon, et al.
Published: (2026)
LLMs as verification oracles for Solidity
by: Bartoletti, Massimo, et al.
Published: (2025)
by: Bartoletti, Massimo, et al.
Published: (2025)
Improving Dynamic Specification Inference with LLM-Generated Counterexamples
by: Balestra, Agustín, et al.
Published: (2026)
by: Balestra, Agustín, et al.
Published: (2026)
Boosting LLMs for Mutation Generation
by: Wang, Bo, et al.
Published: (2026)
by: Wang, Bo, et al.
Published: (2026)
Latent Mutants: A large-scale study on the Interplay between mutation testing and software evolution
by: Sohn, Jeongju, et al.
Published: (2025)
by: Sohn, Jeongju, et al.
Published: (2025)
Software Fairness: An Analysis and Survey
by: Soremekun, Ezekiel, et al.
Published: (2022)
by: Soremekun, Ezekiel, et al.
Published: (2022)
Large-scale, Independent and Comprehensive study of the power of LLMs for test case generation
by: Ouédraogo, Wendkûuni C., et al.
Published: (2024)
by: Ouédraogo, Wendkûuni C., et al.
Published: (2024)
You Can REST Now: Automated REST API Documentation and Testing via LLM-Assisted Request Mutations
by: Decrop, Alix, et al.
Published: (2024)
by: Decrop, Alix, et al.
Published: (2024)
Generating executable oracles to check conformance of client code to requirements of JDK Javadocs using LLMs
by: Jiang, Shan, et al.
Published: (2024)
by: Jiang, Shan, et al.
Published: (2024)
An approach for performance requirements verification and test environments generation
by: Abdeen, Waleed, et al.
Published: (2024)
by: Abdeen, Waleed, et al.
Published: (2024)
Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis
by: Akli, Amal, et al.
Published: (2026)
by: Akli, Amal, et al.
Published: (2026)
Do Code LLMs Do Static Analysis?
by: Su, Chia-Yi, et al.
Published: (2025)
by: Su, Chia-Yi, et al.
Published: (2025)
Observation-based unit test generation at Meta
by: Alshahwan, Nadia, et al.
Published: (2024)
by: Alshahwan, Nadia, et al.
Published: (2024)
Development of an automatic modification system for generated programs using ChatGPT
by: Yoshida, Jun, et al.
Published: (2024)
by: Yoshida, Jun, et al.
Published: (2024)
LLMShot: Reducing snapshot testing maintenance via LLMs
by: Kaynak, Ergün Batuhan, et al.
Published: (2025)
by: Kaynak, Ergün Batuhan, et al.
Published: (2025)
One Model, Many Skills: Parameter-Efficient Fine-Tuning for Multitask Code Analysis
by: Akli, Amal, et al.
Published: (2026)
by: Akli, Amal, et al.
Published: (2026)
Artificial or Just Artful? Do LLMs Bend the Rules in Programming?
by: Sghaier, Oussama Ben, et al.
Published: (2025)
by: Sghaier, Oussama Ben, et al.
Published: (2025)
Development of a Real-Time Simulator Using EMTP-ATP Foreign models for Testing Relays
by: Fabian, Renzo, et al.
Published: (2024)
by: Fabian, Renzo, et al.
Published: (2024)
Summary-Mediated Repair: Can LLMs use code summarisation as a tool for program repair?
by: Twist, Lukas
Published: (2025)
by: Twist, Lukas
Published: (2025)
A11YN: aligning LLMs for accessible web UI code generation
by: Yoon, Janghan, et al.
Published: (2025)
by: Yoon, Janghan, et al.
Published: (2025)
GenCode: A Generic Data Augmentation Framework for Boosting Deep Learning-Based Code Understanding
by: Dong, Zeming, et al.
Published: (2024)
by: Dong, Zeming, et al.
Published: (2024)
When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions
by: Larbi, Maya, et al.
Published: (2025)
by: Larbi, Maya, et al.
Published: (2025)
Learning test generators for cyber-physical systems
by: Peltomäki, Jarkko, et al.
Published: (2024)
by: Peltomäki, Jarkko, et al.
Published: (2024)
OpenAI for OpenAPI: Automated generation of REST API specification via LLMs
by: Chen, Hao, et al.
Published: (2026)
by: Chen, Hao, et al.
Published: (2026)
Test code generation at Ericsson using Program Analysis Augmented Fine Tuned LLMs
by: Krishna, Sai, et al.
Published: (2025)
by: Krishna, Sai, et al.
Published: (2025)
Learning Generalizable Multimodal Representations for Software Vulnerability Detection
by: Dong, Zeming, et al.
Published: (2026)
by: Dong, Zeming, et al.
Published: (2026)
What Types of Code Review Comments Do Developers Most Frequently Resolve?
by: Goldman, Saul, et al.
Published: (2025)
by: Goldman, Saul, et al.
Published: (2025)
Large Language Models for Code Analysis: Do LLMs Really Do Their Job?
by: Fang, Chongzhou, et al.
Published: (2023)
by: Fang, Chongzhou, et al.
Published: (2023)
Boosting Source Code Learning with Text-Oriented Data Augmentation: An Empirical Study
by: Dong, Zeming, et al.
Published: (2023)
by: Dong, Zeming, et al.
Published: (2023)
Do Code LLMs Understand Design Patterns?
by: Pan, Zhenyu, et al.
Published: (2025)
by: Pan, Zhenyu, et al.
Published: (2025)
Quality attributes of test cases and test suites -- importance & challenges from practitioners' perspectives
by: Tran, Huynh Khanh Vi, et al.
Published: (2025)
by: Tran, Huynh Khanh Vi, et al.
Published: (2025)
Do LLMs Provide Links to Code Similar to what they Generate? A Study with Gemini and Bing CoPilot
by: Bifolco, Daniele, et al.
Published: (2025)
by: Bifolco, Daniele, et al.
Published: (2025)
A Comprehensive Study on Large Language Models for Mutation Testing
by: Wang, Bo, et al.
Published: (2024)
by: Wang, Bo, et al.
Published: (2024)
Why Do You Contribute to Stack Overflow? Understanding Cross-Cultural Motivations and Usage Patterns before the Age of LLMs
by: Licorish, Sherlock A., et al.
Published: (2026)
by: Licorish, Sherlock A., et al.
Published: (2026)
To Do or Not to Do: Semantics and Patterns for Do Activities in UML PSSM State Machines
by: Elekes, Márton, et al.
Published: (2023)
by: Elekes, Márton, et al.
Published: (2023)
GenAI-based test case generation and execution in SDV platform
by: Zyberaj, Denesa, et al.
Published: (2025)
by: Zyberaj, Denesa, et al.
Published: (2025)
Similar Items
-
How well LLM-based test generation techniques perform with newer LLM versions?
by: Konstantinou, Michael, et al.
Published: (2026) -
YATE: The Role of Test Repair in LLM-Based Unit Test Generation
by: Konstantinou, Michael, et al.
Published: (2025) -
Bounded Synthesis of Synchronized Distributed Models from Lightweight Specifications
by: Castro, Pablo F., et al.
Published: (2025) -
Evaluating LLMs for One-Shot Patching of Real and Artificial Vulnerabilities
by: Garg, Aayush, et al.
Published: (2025) -
Inferring Code Correctness from Specification
by: Florian, Tambon, et al.
Published: (2026)