Consistency Meets Verification: Enhancing Test Generation Quality in Large Language Models Without Ground-Truth Solutions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Taherkhani, Hamed, DaghighFarsoodeh, Alireza, Chowdhury, Mohammad, Pham, Hung Viet, Hemmati, Hadi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Task-oriented Prompt Enhancement via Script Generation
von: Wang, Chung-Yu, et al.
Veröffentlicht: (2024)
von: Wang, Chung-Yu, et al.
Veröffentlicht: (2024)
Selection of Prompt Engineering Techniques for Code Generation through Predicting Code Complexity
von: Wang, Chung-Yu, et al.
Veröffentlicht: (2024)
von: Wang, Chung-Yu, et al.
Veröffentlicht: (2024)
RGFL: Reasoning Guided Fault Localization for Automated Program Repair Using Large Language Models
von: Sepidband, Melika, et al.
Veröffentlicht: (2026)
von: Sepidband, Melika, et al.
Veröffentlicht: (2026)
Automated Prompt Engineering for Cost-Effective Code Generation Using Evolutionary Algorithm
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2024)
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2024)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
von: Daghighfarsoodeh, Alireza, et al.
Veröffentlicht: (2025)
von: Daghighfarsoodeh, Alireza, et al.
Veröffentlicht: (2025)
Enhancing LLM-Based Code Generation with Complexity Metrics: A Feedback-Driven Approach
von: Sepidband, Melika, et al.
Veröffentlicht: (2025)
von: Sepidband, Melika, et al.
Veröffentlicht: (2025)
On the Role of Fault Localization Context for LLM-Based Program Repair
von: Sepidband, Melika, et al.
Veröffentlicht: (2026)
von: Sepidband, Melika, et al.
Veröffentlicht: (2026)
Can ChatGPT Support Developers? An Empirical Evaluation of Large Language Models for Code Generation
von: Jin, Kailun, et al.
Veröffentlicht: (2024)
von: Jin, Kailun, et al.
Veröffentlicht: (2024)
Toward Automated Validation of Language Model Synthesized Test Cases using Semantic Entropy
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2024)
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2024)
FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair
von: Fatima, Sakina, et al.
Veröffentlicht: (2023)
von: Fatima, Sakina, et al.
Veröffentlicht: (2023)
A Systematic Mapping Study of Crowd Knowledge Enhanced Software Engineering Research Using Stack Overflow
von: Tanzil, Minaoar, et al.
Veröffentlicht: (2024)
von: Tanzil, Minaoar, et al.
Veröffentlicht: (2024)
An Empirical Study on Bug Severity Estimation using Source Code Metrics and Static Analysis
von: Mashhadi, Ehsan, et al.
Veröffentlicht: (2022)
von: Mashhadi, Ehsan, et al.
Veröffentlicht: (2022)
Assessing Evaluation Metrics for Neural Test Oracle Generation
von: Shin, Jiho, et al.
Veröffentlicht: (2023)
von: Shin, Jiho, et al.
Veröffentlicht: (2023)
Domain Adaptation for Code Model-based Unit Test Case Generation
von: Shin, Jiho, et al.
Veröffentlicht: (2023)
von: Shin, Jiho, et al.
Veröffentlicht: (2023)
Program Slicing in the Era of Large Language Models
von: Shahandashti, Kimya Khakzad, et al.
Veröffentlicht: (2024)
von: Shahandashti, Kimya Khakzad, et al.
Veröffentlicht: (2024)
ZeroCoder: Can LLMs Improve Code Generation Without Ground-Truth Supervision?
von: Fan, Lishui, et al.
Veröffentlicht: (2026)
von: Fan, Lishui, et al.
Veröffentlicht: (2026)
StaAgent: An Agentic Framework for Testing Static Analyzers
von: Nnorom, Elijah, et al.
Veröffentlicht: (2025)
von: Nnorom, Elijah, et al.
Veröffentlicht: (2025)
ABTest: Behavior-Driven Testing for AI Coding Agents
von: Dai, Wuyang, et al.
Veröffentlicht: (2026)
von: Dai, Wuyang, et al.
Veröffentlicht: (2026)
Applications and Challenges of Fairness APIs in Machine Learning Software
von: Das, Ajoy, et al.
Veröffentlicht: (2025)
von: Das, Ajoy, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Test Generation: How Far Are We?
von: Shin, Jiho, et al.
Veröffentlicht: (2024)
von: Shin, Jiho, et al.
Veröffentlicht: (2024)
Defect Prediction with Content-based Features
von: Pham, Hung Viet, et al.
Veröffentlicht: (2024)
von: Pham, Hung Viet, et al.
Veröffentlicht: (2024)
GAN-enhanced Simulation-driven DNN Testing in Absence of Ground Truth
von: Attaoui, Mohammed, et al.
Veröffentlicht: (2025)
von: Attaoui, Mohammed, et al.
Veröffentlicht: (2025)
Automatic Instantiation of Assurance Cases from Patterns Using Large Language Models
von: Odu, Oluwafemi, et al.
Veröffentlicht: (2024)
von: Odu, Oluwafemi, et al.
Veröffentlicht: (2024)
Demystifying Errors in LLM Reasoning Traces: An Empirical Study of Code Execution Simulation
von: Abdollahi, Mohammad, et al.
Veröffentlicht: (2025)
von: Abdollahi, Mohammad, et al.
Veröffentlicht: (2025)
Detecting Call Graph Unsoundness without Ground Truth
von: Zhong, Fangtian, et al.
Veröffentlicht: (2026)
von: Zhong, Fangtian, et al.
Veröffentlicht: (2026)
Revisiting Code Debloating with Ground Truth-based Evaluation
von: Bilal, Muhammad, et al.
Veröffentlicht: (2026)
von: Bilal, Muhammad, et al.
Veröffentlicht: (2026)
DocTer: Documentation Guided Fuzzing for Testing Deep Learning API Functions
von: Xie, Danning, et al.
Veröffentlicht: (2021)
von: Xie, Danning, et al.
Veröffentlicht: (2021)
Prompt Engineering or Fine-Tuning: An Empirical Assessment of LLMs for Code
von: Shin, Jiho, et al.
Veröffentlicht: (2023)
von: Shin, Jiho, et al.
Veröffentlicht: (2023)
Engineering Pitfalls in AI Coding Tools: An Empirical Study of Bugs in Claude Code, Codex, and Gemini CLI
von: Zhang, Ruixin, et al.
Veröffentlicht: (2026)
von: Zhang, Ruixin, et al.
Veröffentlicht: (2026)
Formal Verification of Consistency for Systems with Redundant Controllers
von: Johansson, Bjarne, et al.
Veröffentlicht: (2024)
von: Johansson, Bjarne, et al.
Veröffentlicht: (2024)
ENCORE: Ensemble Learning using Convolution Neural Machine Translation for Automatic Program Repair
von: Lutellier, Thibaud, et al.
Veröffentlicht: (2019)
von: Lutellier, Thibaud, et al.
Veröffentlicht: (2019)
SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection
von: Ahmed, Md Basim Uddin, et al.
Veröffentlicht: (2025)
von: Ahmed, Md Basim Uddin, et al.
Veröffentlicht: (2025)
A General Solution for the Implementation of CI/CD in Embedded Linux Development
von: Agahi, Behnam, et al.
Veröffentlicht: (2025)
von: Agahi, Behnam, et al.
Veröffentlicht: (2025)
Quality Assessment of Python Tests Generated by Large Language Models
von: Alves, Victor, et al.
Veröffentlicht: (2025)
von: Alves, Victor, et al.
Veröffentlicht: (2025)
The Rise of Agentic Testing: Multi-Agent Systems for Robust Software Quality Assurance
von: Naqvi, Saba, et al.
Veröffentlicht: (2026)
von: Naqvi, Saba, et al.
Veröffentlicht: (2026)
Combining Tests and Proofs for Better Software Verification
von: Huang, Li, et al.
Veröffentlicht: (2026)
von: Huang, Li, et al.
Veröffentlicht: (2026)
A Ground-Truth-Based Evaluation of Vulnerability Detection Across Multiple Ecosystems
von: Mandl, Peter, et al.
Veröffentlicht: (2026)
von: Mandl, Peter, et al.
Veröffentlicht: (2026)
Watchdogs and Oracles: Runtime Verification Meets Large Language Models for Autonomous Systems
von: Ferrando, Angelo
Veröffentlicht: (2025)
von: Ferrando, Angelo
Veröffentlicht: (2025)
Checkification: A Practical Approach for Testing Static Analysis Truths
von: Ferreiro, Daniela, et al.
Veröffentlicht: (2025)
von: Ferreiro, Daniela, et al.
Veröffentlicht: (2025)
A Systematic Approach for Assessing Large Language Models' Test Case Generation Capability
von: Chang, Hung-Fu, et al.
Veröffentlicht: (2025)
von: Chang, Hung-Fu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Task-oriented Prompt Enhancement via Script Generation
von: Wang, Chung-Yu, et al.
Veröffentlicht: (2024) -
Selection of Prompt Engineering Techniques for Code Generation through Predicting Code Complexity
von: Wang, Chung-Yu, et al.
Veröffentlicht: (2024) -
RGFL: Reasoning Guided Fault Localization for Automated Program Repair Using Large Language Models
von: Sepidband, Melika, et al.
Veröffentlicht: (2026) -
Automated Prompt Engineering for Cost-Effective Code Generation Using Evolutionary Algorithm
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2024) -
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
von: Daghighfarsoodeh, Alireza, et al.
Veröffentlicht: (2025)