Assessing Evaluation Metrics for Neural Test Oracle Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shin, Jiho, Hemmati, Hadi, Wei, Moshi, Wang, Song |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Domain Adaptation for Code Model-based Unit Test Case Generation
von: Shin, Jiho, et al.
Veröffentlicht: (2023)
von: Shin, Jiho, et al.
Veröffentlicht: (2023)
The Good, the Bad, and the Missing: Neural Code Generation for Machine Learning Tasks
von: Shin, Jiho, et al.
Veröffentlicht: (2023)
von: Shin, Jiho, et al.
Veröffentlicht: (2023)
Retrieval-Augmented Test Generation: How Far Are We?
von: Shin, Jiho, et al.
Veröffentlicht: (2024)
von: Shin, Jiho, et al.
Veröffentlicht: (2024)
Prompt Engineering or Fine-Tuning: An Empirical Assessment of LLMs for Code
von: Shin, Jiho, et al.
Veröffentlicht: (2023)
von: Shin, Jiho, et al.
Veröffentlicht: (2023)
Nexus: Execution-Grounded Multi-Agent Test Oracle Synthesis
von: Huang, Dong, et al.
Veröffentlicht: (2025)
von: Huang, Dong, et al.
Veröffentlicht: (2025)
Enhancing LLM-Based Code Generation with Complexity Metrics: A Feedback-Driven Approach
von: Sepidband, Melika, et al.
Veröffentlicht: (2025)
von: Sepidband, Melika, et al.
Veröffentlicht: (2025)
Evaluating LLMs on Sequential API Call Through Automated Test Generation
von: Huang, Yuheng, et al.
Veröffentlicht: (2025)
von: Huang, Yuheng, et al.
Veröffentlicht: (2025)
Toward Automated Validation of Language Model Synthesized Test Cases using Semantic Entropy
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2024)
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2024)
Automated Trustworthiness Oracle Generation for Machine Learning Text Classifiers
von: Tung, Lam Nguyen, et al.
Veröffentlicht: (2024)
von: Tung, Lam Nguyen, et al.
Veröffentlicht: (2024)
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
von: Yang, Zheyuan, et al.
Veröffentlicht: (2025)
von: Yang, Zheyuan, et al.
Veröffentlicht: (2025)
GenX: Mastering Code and Test Generation with Execution Feedback
von: Wang, Nan, et al.
Veröffentlicht: (2024)
von: Wang, Nan, et al.
Veröffentlicht: (2024)
CodeContests+: High-Quality Test Case Generation for Competitive Programming
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
SemLink: A Semantic-Aware Automated Test Oracle for Hyperlink Verification using Siamese Sentence-BERT
von: Yang, Guan-Yan, et al.
Veröffentlicht: (2026)
von: Yang, Guan-Yan, et al.
Veröffentlicht: (2026)
Measuring the Influence of Incorrect Code on Test Generation
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
von: Liu, Steven, et al.
Veröffentlicht: (2026)
von: Liu, Steven, et al.
Veröffentlicht: (2026)
FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
An Empirical Study on Bug Severity Estimation using Source Code Metrics and Static Analysis
von: Mashhadi, Ehsan, et al.
Veröffentlicht: (2022)
von: Mashhadi, Ehsan, et al.
Veröffentlicht: (2022)
GUI Test Migration via Abstraction and Concretization
von: Zhang, Yakun, et al.
Veröffentlicht: (2024)
von: Zhang, Yakun, et al.
Veröffentlicht: (2024)
HarnessLLM: Automatic Testing Harness Generation via Reinforcement Learning
von: Liu, Yujian, et al.
Veröffentlicht: (2025)
von: Liu, Yujian, et al.
Veröffentlicht: (2025)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
von: Huang, Dong, et al.
Veröffentlicht: (2025)
von: Huang, Dong, et al.
Veröffentlicht: (2025)
DuET: Dual Execution for Test Output Prediction with Generated Code and Pseudocode
von: Han, Hojae, et al.
Veröffentlicht: (2026)
von: Han, Hojae, et al.
Veröffentlicht: (2026)
Automated Discovery of Test Oracles for Database Management Systems Using LLMs
von: Mang, Qiuyang, et al.
Veröffentlicht: (2025)
von: Mang, Qiuyang, et al.
Veröffentlicht: (2025)
Evaluation of Code LLMs on Geospatial Code Generation
von: Gramacki, Piotr, et al.
Veröffentlicht: (2024)
von: Gramacki, Piotr, et al.
Veröffentlicht: (2024)
UnitCoder: Scalable Iterative Code Synthesis with Unit Test Guidance
von: Ma, Yichuan, et al.
Veröffentlicht: (2025)
von: Ma, Yichuan, et al.
Veröffentlicht: (2025)
Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval
von: Wang, Jiexin, et al.
Veröffentlicht: (2024)
von: Wang, Jiexin, et al.
Veröffentlicht: (2024)
Checker Bug Detection and Repair in Deep Learning Libraries
von: Harzevili, Nima Shiri, et al.
Veröffentlicht: (2024)
von: Harzevili, Nima Shiri, et al.
Veröffentlicht: (2024)
DiffuTester: Accelerating Unit Test Generation for Diffusion LLMs via Mining Structural Pattern
von: Yang, Lekang, et al.
Veröffentlicht: (2025)
von: Yang, Lekang, et al.
Veröffentlicht: (2025)
Automated Prompt Engineering for Cost-Effective Code Generation Using Evolutionary Algorithm
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2024)
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2024)
Measuring Complexity at the Requirements Stage: Spectral Metrics as Development Effort Predictors
von: Vierlboeck, Maximilian, et al.
Veröffentlicht: (2026)
von: Vierlboeck, Maximilian, et al.
Veröffentlicht: (2026)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
von: Du, Yongkang, et al.
Veröffentlicht: (2025)
von: Du, Yongkang, et al.
Veröffentlicht: (2025)
Using Large Language Models for Student-Code Guided Test Case Generation in Computer Science Education
von: Kumar, Nischal Ashok, et al.
Veröffentlicht: (2024)
von: Kumar, Nischal Ashok, et al.
Veröffentlicht: (2024)
Testing and Evaluation of Large Language Models: Correctness, Non-Toxicity, and Fairness
von: Wang, Wenxuan
Veröffentlicht: (2024)
von: Wang, Wenxuan
Veröffentlicht: (2024)
Consistency Meets Verification: Enhancing Test Generation Quality in Large Language Models Without Ground-Truth Solutions
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2026)
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2026)
Evaluating Language Models for Efficient Code Generation
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
HateModerate: Testing Hate Speech Detectors against Content Moderation Policies
von: Zheng, Jiangrui, et al.
Veröffentlicht: (2023)
von: Zheng, Jiangrui, et al.
Veröffentlicht: (2023)
Machine Translation Testing via Syntactic Tree Pruning
von: Zhang, Quanjun, et al.
Veröffentlicht: (2024)
von: Zhang, Quanjun, et al.
Veröffentlicht: (2024)
StaAgent: An Agentic Framework for Testing Static Analyzers
von: Nnorom, Elijah, et al.
Veröffentlicht: (2025)
von: Nnorom, Elijah, et al.
Veröffentlicht: (2025)
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
von: Zhang, Chenchen, et al.
Veröffentlicht: (2025)
von: Zhang, Chenchen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Domain Adaptation for Code Model-based Unit Test Case Generation
von: Shin, Jiho, et al.
Veröffentlicht: (2023) -
The Good, the Bad, and the Missing: Neural Code Generation for Machine Learning Tasks
von: Shin, Jiho, et al.
Veröffentlicht: (2023) -
Retrieval-Augmented Test Generation: How Far Are We?
von: Shin, Jiho, et al.
Veröffentlicht: (2024) -
Prompt Engineering or Fine-Tuning: An Empirical Assessment of LLMs for Code
von: Shin, Jiho, et al.
Veröffentlicht: (2023) -
Nexus: Execution-Grounded Multi-Agent Test Oracle Synthesis
von: Huang, Dong, et al.
Veröffentlicht: (2025)