Salvato in:
| Autori principali: | Zhao, Fufangchen, Jin, Guoqiang, Zhao, Rui, Huang, Jiangheng, Tan, Fei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2407.17150 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Consistency Matters: Explore LLMs Consistency From a Black-Box Perspective
di: Zhao, Fufangchen, et al.
Pubblicazione: (2024)
di: Zhao, Fufangchen, et al.
Pubblicazione: (2024)
Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study
di: Li, Bowen, et al.
Pubblicazione: (2024)
di: Li, Bowen, et al.
Pubblicazione: (2024)
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
di: Yang, Zheyuan, et al.
Pubblicazione: (2025)
di: Yang, Zheyuan, et al.
Pubblicazione: (2025)
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
di: Liu, Steven, et al.
Pubblicazione: (2026)
di: Liu, Steven, et al.
Pubblicazione: (2026)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
di: Du, Yongkang, et al.
Pubblicazione: (2025)
di: Du, Yongkang, et al.
Pubblicazione: (2025)
Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability
di: Chen, Zejian, et al.
Pubblicazione: (2026)
di: Chen, Zejian, et al.
Pubblicazione: (2026)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
di: Huang, Dong, et al.
Pubblicazione: (2025)
di: Huang, Dong, et al.
Pubblicazione: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
di: Li, Jia, et al.
Pubblicazione: (2024)
di: Li, Jia, et al.
Pubblicazione: (2024)
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
di: Wang, Yibo, et al.
Pubblicazione: (2025)
di: Wang, Yibo, et al.
Pubblicazione: (2025)
Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations
di: Salimian, Sina, et al.
Pubblicazione: (2025)
di: Salimian, Sina, et al.
Pubblicazione: (2025)
Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems
di: Cao, Yuhan, et al.
Pubblicazione: (2025)
di: Cao, Yuhan, et al.
Pubblicazione: (2025)
A Comparative Study on the Impact of Test-Driven Development (TDD) and Behavior-Driven Development (BDD) on Enterprise Software Delivery Effectiveness
di: Cui, Jun
Pubblicazione: (2024)
di: Cui, Jun
Pubblicazione: (2024)
AutoTestForge: A Multidimensional Automated Testing Framework for Natural Language Processing Models
di: Xing, Hengrui, et al.
Pubblicazione: (2025)
di: Xing, Hengrui, et al.
Pubblicazione: (2025)
MathDuels: Evaluating LLMs as Problem Posers and Solvers
di: Xu, Zhiqiu, et al.
Pubblicazione: (2026)
di: Xu, Zhiqiu, et al.
Pubblicazione: (2026)
DiffuTester: Accelerating Unit Test Generation for Diffusion LLMs via Mining Structural Pattern
di: Yang, Lekang, et al.
Pubblicazione: (2025)
di: Yang, Lekang, et al.
Pubblicazione: (2025)
Evaluating LLMs on Sequential API Call Through Automated Test Generation
di: Huang, Yuheng, et al.
Pubblicazione: (2025)
di: Huang, Yuheng, et al.
Pubblicazione: (2025)
Showing LLM-Generated Code Selectively Based on Confidence of LLMs
di: Li, Jia, et al.
Pubblicazione: (2024)
di: Li, Jia, et al.
Pubblicazione: (2024)
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
di: Tian, Yuchen, et al.
Pubblicazione: (2024)
di: Tian, Yuchen, et al.
Pubblicazione: (2024)
GUI Test Migration via Abstraction and Concretization
di: Zhang, Yakun, et al.
Pubblicazione: (2024)
di: Zhang, Yakun, et al.
Pubblicazione: (2024)
Multi-Programming Language Sandbox for LLMs
di: Dou, Shihan, et al.
Pubblicazione: (2024)
di: Dou, Shihan, et al.
Pubblicazione: (2024)
RePair: Automated Program Repair with Process-based Feedback
di: Zhao, Yuze, et al.
Pubblicazione: (2024)
di: Zhao, Yuze, et al.
Pubblicazione: (2024)
Measuring the Influence of Incorrect Code on Test Generation
di: Huang, Dong, et al.
Pubblicazione: (2024)
di: Huang, Dong, et al.
Pubblicazione: (2024)
Unlocking a New Rust Programming Experience: Fast and Slow Thinking with LLMs to Conquer Undefined Behaviors
di: Jiang, Renshuang, et al.
Pubblicazione: (2025)
di: Jiang, Renshuang, et al.
Pubblicazione: (2025)
DuET: Dual Execution for Test Output Prediction with Generated Code and Pseudocode
di: Han, Hojae, et al.
Pubblicazione: (2026)
di: Han, Hojae, et al.
Pubblicazione: (2026)
Solver-Independent Automated Problem Formulation via LLMs for High-Cost Simulation-Driven Design
di: Li, Yuchen, et al.
Pubblicazione: (2025)
di: Li, Yuchen, et al.
Pubblicazione: (2025)
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
di: Li, Jia, et al.
Pubblicazione: (2024)
di: Li, Jia, et al.
Pubblicazione: (2024)
OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs
di: Ahmad, Wasi Uddin, et al.
Pubblicazione: (2025)
di: Ahmad, Wasi Uddin, et al.
Pubblicazione: (2025)
DependEval: Benchmarking LLMs for Repository Dependency Understanding
di: Du, Junjia, et al.
Pubblicazione: (2025)
di: Du, Junjia, et al.
Pubblicazione: (2025)
From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future
di: Jin, Haolin, et al.
Pubblicazione: (2024)
di: Jin, Haolin, et al.
Pubblicazione: (2024)
CodePivot: Bootstrapping Multilingual Transpilation in LLMs via Reinforcement Learning without Parallel Corpora
di: Li, Shangyu, et al.
Pubblicazione: (2026)
di: Li, Shangyu, et al.
Pubblicazione: (2026)
Nexus: Execution-Grounded Multi-Agent Test Oracle Synthesis
di: Huang, Dong, et al.
Pubblicazione: (2025)
di: Huang, Dong, et al.
Pubblicazione: (2025)
AC4: Algebraic Computation Checker for Circuit Constraints in ZKPs
di: Yang, Qizhe, et al.
Pubblicazione: (2024)
di: Yang, Qizhe, et al.
Pubblicazione: (2024)
CCT-Code: Cross-Consistency Training for Multilingual Clone Detection and Code Search
di: Tikhonov, Anton, et al.
Pubblicazione: (2023)
di: Tikhonov, Anton, et al.
Pubblicazione: (2023)
Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
di: Li, Zhuohao, et al.
Pubblicazione: (2025)
di: Li, Zhuohao, et al.
Pubblicazione: (2025)
ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs
di: Ding, Peng, et al.
Pubblicazione: (2025)
di: Ding, Peng, et al.
Pubblicazione: (2025)
ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
di: Ding, Xianzhong, et al.
Pubblicazione: (2026)
di: Ding, Xianzhong, et al.
Pubblicazione: (2026)
Alibaba LingmaAgent: Improving Automated Issue Resolution via Comprehensive Repository Exploration
di: Ma, Yingwei, et al.
Pubblicazione: (2024)
di: Ma, Yingwei, et al.
Pubblicazione: (2024)
CodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case Generation
di: Cai, Jianfeng, et al.
Pubblicazione: (2026)
di: Cai, Jianfeng, et al.
Pubblicazione: (2026)
TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
di: Ahmed, Toufique, et al.
Pubblicazione: (2024)
di: Ahmed, Toufique, et al.
Pubblicazione: (2024)
Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey
di: Wang, Junqiao, et al.
Pubblicazione: (2024)
di: Wang, Junqiao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Consistency Matters: Explore LLMs Consistency From a Black-Box Perspective
di: Zhao, Fufangchen, et al.
Pubblicazione: (2024) -
Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study
di: Li, Bowen, et al.
Pubblicazione: (2024) -
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
di: Yang, Zheyuan, et al.
Pubblicazione: (2025) -
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
di: Liu, Steven, et al.
Pubblicazione: (2026) -
FairCoder: Evaluating Social Bias of LLMs in Code Generation
di: Du, Yongkang, et al.
Pubblicazione: (2025)