Measuring the Influence of Incorrect Code on Test Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Huang, Dong, Zhang, Jie M., Harman, Mark, Du, Mingzhe, Cui, Heming |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Benchmarking LLMs for Unit Test Generation from Real-World Functions
par: Huang, Dong, et autres
Publié: (2025)
par: Huang, Dong, et autres
Publié: (2025)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
par: Huang, Dong, et autres
Publié: (2024)
par: Huang, Dong, et autres
Publié: (2024)
EffiLearner: Enhancing Efficiency of Generated Code via Self-Optimization
par: Huang, Dong, et autres
Publié: (2024)
par: Huang, Dong, et autres
Publié: (2024)
EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning
par: Huang, Dong, et autres
Publié: (2024)
par: Huang, Dong, et autres
Publié: (2024)
Nexus: Execution-Grounded Multi-Agent Test Oracle Synthesis
par: Huang, Dong, et autres
Publié: (2025)
par: Huang, Dong, et autres
Publié: (2025)
Mercury: A Code Efficiency Benchmark for Code Large Language Models
par: Du, Mingzhe, et autres
Publié: (2024)
par: Du, Mingzhe, et autres
Publié: (2024)
Bias Testing and Mitigation in LLM-based Code Generation
par: Huang, Dong, et autres
Publié: (2023)
par: Huang, Dong, et autres
Publié: (2023)
Dynamic Scaling of Unit Tests for Code Reward Modeling
par: Ma, Zeyao, et autres
Publié: (2025)
par: Ma, Zeyao, et autres
Publié: (2025)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
par: Du, Yongkang, et autres
Publié: (2025)
par: Du, Yongkang, et autres
Publié: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
par: Liu, Steven, et autres
Publié: (2026)
par: Liu, Steven, et autres
Publié: (2026)
GenX: Mastering Code and Test Generation with Execution Feedback
par: Wang, Nan, et autres
Publié: (2024)
par: Wang, Nan, et autres
Publié: (2024)
An Empirical Study of the Non-determinism of ChatGPT in Code Generation
par: Ouyang, Shuyin, et autres
Publié: (2023)
par: Ouyang, Shuyin, et autres
Publié: (2023)
Measuring LLM Code Generation Stability via Structural Entropy
par: Song, Yewei, et autres
Publié: (2025)
par: Song, Yewei, et autres
Publié: (2025)
CodeContests+: High-Quality Test Case Generation for Competitive Programming
par: Wang, Zihan, et autres
Publié: (2025)
par: Wang, Zihan, et autres
Publié: (2025)
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning
par: Wang, Hao, et autres
Publié: (2026)
par: Wang, Hao, et autres
Publié: (2026)
YATE: The Role of Test Repair in LLM-Based Unit Test Generation
par: Konstantinou, Michael, et autres
Publié: (2025)
par: Konstantinou, Michael, et autres
Publié: (2025)
CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation
par: Huang, Dong, et autres
Publié: (2023)
par: Huang, Dong, et autres
Publié: (2023)
DuET: Dual Execution for Test Output Prediction with Generated Code and Pseudocode
par: Han, Hojae, et autres
Publié: (2026)
par: Han, Hojae, et autres
Publié: (2026)
Towards Exception Safety Code Generation with Intermediate Representation Agents Framework
par: Zhang, Xuanming, et autres
Publié: (2024)
par: Zhang, Xuanming, et autres
Publié: (2024)
Is Your AI-Generated Code Really Safe? Evaluating Large Language Models on Secure Code Generation with CodeSecEval
par: Wang, Jiexin, et autres
Publié: (2024)
par: Wang, Jiexin, et autres
Publié: (2024)
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
Code Fingerprints: Disentangled Attribution of LLM-Generated Code
par: Guo, Jiaxun, et autres
Publié: (2026)
par: Guo, Jiaxun, et autres
Publié: (2026)
Seeker: Towards Exception Safety Code Generation with Intermediate Language Agents Framework
par: Zhang, Xuanming, et autres
Publié: (2024)
par: Zhang, Xuanming, et autres
Publié: (2024)
DeCon: Detecting Incorrect Assertions via Postconditions Generated by a Large Language Model
par: Yu, Hao, et autres
Publié: (2025)
par: Yu, Hao, et autres
Publié: (2025)
Evaluating and Achieving Controllable Code Completion in Code LLM
par: Zhang, Jiajun, et autres
Publié: (2026)
par: Zhang, Jiajun, et autres
Publié: (2026)
RethinkMCTS: Refining Erroneous Thoughts in Monte Carlo Tree Search for Code Generation
par: Li, Qingyao, et autres
Publié: (2024)
par: Li, Qingyao, et autres
Publié: (2024)
Environment-Aware Code Generation: How far are We?
par: Wu, Tongtong, et autres
Publié: (2026)
par: Wu, Tongtong, et autres
Publié: (2026)
Using Large Language Models for Student-Code Guided Test Case Generation in Computer Science Education
par: Kumar, Nischal Ashok, et autres
Publié: (2024)
par: Kumar, Nischal Ashok, et autres
Publié: (2024)
Evaluation of Code LLMs on Geospatial Code Generation
par: Gramacki, Piotr, et autres
Publié: (2024)
par: Gramacki, Piotr, et autres
Publié: (2024)
Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey
par: Wang, Junqiao, et autres
Publié: (2024)
par: Wang, Junqiao, et autres
Publié: (2024)
Iterative Refinement of Project-Level Code Context for Precise Code Generation with Compiler Feedback
par: Bi, Zhangqian, et autres
Publié: (2024)
par: Bi, Zhangqian, et autres
Publié: (2024)
The Role of DevOps in Enhancing Enterprise Software Delivery Success through R&D Efficiency and Source Code Management
par: Cui, Jun
Publié: (2024)
par: Cui, Jun
Publié: (2024)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
par: Chou, Jason, et autres
Publié: (2025)
par: Chou, Jason, et autres
Publié: (2025)
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
par: Wang, Yibo, et autres
Publié: (2025)
par: Wang, Yibo, et autres
Publié: (2025)
Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasks
par: Zhao, Songwen, et autres
Publié: (2025)
par: Zhao, Songwen, et autres
Publié: (2025)
VersiCode: Towards Version-controllable Code Generation
par: Wu, Tongtong, et autres
Publié: (2024)
par: Wu, Tongtong, et autres
Publié: (2024)
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
CWEval: Outcome-driven Evaluation on Functionality and Security of LLM Code Generation
par: Peng, Jinjun, et autres
Publié: (2025)
par: Peng, Jinjun, et autres
Publié: (2025)
A Comparative Study on the Impact of Test-Driven Development (TDD) and Behavior-Driven Development (BDD) on Enterprise Software Delivery Effectiveness
par: Cui, Jun
Publié: (2024)
par: Cui, Jun
Publié: (2024)
Documents similaires
-
Benchmarking LLMs for Unit Test Generation from Real-World Functions
par: Huang, Dong, et autres
Publié: (2025) -
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
par: Huang, Dong, et autres
Publié: (2024) -
EffiLearner: Enhancing Efficiency of Generated Code via Self-Optimization
par: Huang, Dong, et autres
Publié: (2024) -
EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuning
par: Huang, Dong, et autres
Publié: (2024) -
Nexus: Execution-Grounded Multi-Agent Test Oracle Synthesis
par: Huang, Dong, et autres
Publié: (2025)