HarnessLLM: Automatic Testing Harness Generation via Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yujian, Ji, Jiabao, Zhang, Yang, Guo, Wenbo, Jaakkola, Tommi, Chang, Shiyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
by: Lin, Jiahang, et al.
Published: (2026)
by: Lin, Jiahang, et al.
Published: (2026)
Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning
by: Liu, Yujian, et al.
Published: (2024)
by: Liu, Yujian, et al.
Published: (2024)
How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
by: Liu, Yujian, et al.
Published: (2026)
by: Liu, Yujian, et al.
Published: (2026)
Absorber LLM: Harnessing Causal Synchronization for Test-Time Training
by: Zhang, Zhixin, et al.
Published: (2026)
by: Zhang, Zhixin, et al.
Published: (2026)
Revisiting Who's Harry Potter: Towards Targeted Unlearning from a Causal Intervention Perspective
by: Liu, Yujian, et al.
Published: (2024)
by: Liu, Yujian, et al.
Published: (2024)
HarnessAgent: Scaling Automatic Fuzzing Harness Construction with Tool-Augmented LLM Pipelines
by: Yang, Kang, et al.
Published: (2025)
by: Yang, Kang, et al.
Published: (2025)
DiffuTester: Accelerating Unit Test Generation for Diffusion LLMs via Mining Structural Pattern
by: Yang, Lekang, et al.
Published: (2025)
by: Yang, Lekang, et al.
Published: (2025)
Effective Harness Engineering for Algorithm Discovery with Coding Agents
by: Ishibashi, Yoichi, et al.
Published: (2026)
by: Ishibashi, Yoichi, et al.
Published: (2026)
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
by: Liu, Steven, et al.
Published: (2026)
by: Liu, Steven, et al.
Published: (2026)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
Code Fingerprints: Disentangled Attribution of LLM-Generated Code
by: Guo, Jiaxun, et al.
Published: (2026)
by: Guo, Jiaxun, et al.
Published: (2026)
StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback
by: Dou, Shihan, et al.
Published: (2024)
by: Dou, Shihan, et al.
Published: (2024)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
by: Huang, Dong, et al.
Published: (2024)
by: Huang, Dong, et al.
Published: (2024)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
by: Chou, Jason, et al.
Published: (2025)
by: Chou, Jason, et al.
Published: (2025)
MORTAR: Multi-turn Metamorphic Testing for LLM-based Dialogue Systems
by: Guo, Guoxiang, et al.
Published: (2024)
by: Guo, Guoxiang, et al.
Published: (2024)
GUI Test Migration via Abstraction and Concretization
by: Zhang, Yakun, et al.
Published: (2024)
by: Zhang, Yakun, et al.
Published: (2024)
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
SAFE: Harnessing LLM for Scenario-Driven ADS Testing from Multimodal Crash Data
by: Luo, Siwei, et al.
Published: (2025)
by: Luo, Siwei, et al.
Published: (2025)
LLMSniffer: Detecting LLM-Generated Code via GraphCodeBERT and Supervised Contrastive Learning
by: Dihan, Mahir Labib, et al.
Published: (2026)
by: Dihan, Mahir Labib, et al.
Published: (2026)
Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey
by: Wang, Junqiao, et al.
Published: (2024)
by: Wang, Junqiao, et al.
Published: (2024)
CodeContests+: High-Quality Test Case Generation for Competitive Programming
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
(Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs
by: Ma, Wanqin, et al.
Published: (2023)
by: Ma, Wanqin, et al.
Published: (2023)
CoCoST: Automatic Complex Code Generation with Online Searching and Correctness Testing
by: He, Xinyi, et al.
Published: (2024)
by: He, Xinyi, et al.
Published: (2024)
Machine Translation Testing via Syntactic Tree Pruning
by: Zhang, Quanjun, et al.
Published: (2024)
by: Zhang, Quanjun, et al.
Published: (2024)
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Measuring LLM Code Generation Stability via Structural Entropy
by: Song, Yewei, et al.
Published: (2025)
by: Song, Yewei, et al.
Published: (2025)
Measuring the Influence of Incorrect Code on Test Generation
by: Huang, Dong, et al.
Published: (2024)
by: Huang, Dong, et al.
Published: (2024)
GenX: Mastering Code and Test Generation with Execution Feedback
by: Wang, Nan, et al.
Published: (2024)
by: Wang, Nan, et al.
Published: (2024)
LLM-as-a-Judge for Reference-less Automatic Code Validation and Refinement for Natural Language to Bash in IT Automation
by: Vo, Ngoc Phuoc An, et al.
Published: (2025)
by: Vo, Ngoc Phuoc An, et al.
Published: (2025)
Optimizing Case-Based Reasoning System for Functional Test Script Generation with Large Language Models
by: Guo, Siyuan, et al.
Published: (2025)
by: Guo, Siyuan, et al.
Published: (2025)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
by: Huang, Dong, et al.
Published: (2025)
by: Huang, Dong, et al.
Published: (2025)
Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback
by: Skopin, Egor, et al.
Published: (2026)
by: Skopin, Egor, et al.
Published: (2026)
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
by: Yang, Zheyuan, et al.
Published: (2025)
by: Yang, Zheyuan, et al.
Published: (2025)
ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
by: Zhang, Chenchen, et al.
Published: (2025)
by: Zhang, Chenchen, et al.
Published: (2025)
Learning to Commit: Generating Organic Pull Requests via Online Repository Memory
by: Li, Mo, et al.
Published: (2026)
by: Li, Mo, et al.
Published: (2026)
Evaluating and Achieving Controllable Code Completion in Code LLM
by: Zhang, Jiajun, et al.
Published: (2026)
by: Zhang, Jiajun, et al.
Published: (2026)
CodePivot: Bootstrapping Multilingual Transpilation in LLMs via Reinforcement Learning without Parallel Corpora
by: Li, Shangyu, et al.
Published: (2026)
by: Li, Shangyu, et al.
Published: (2026)
ProbeLLM: Automating Principled Diagnosis of LLM Failures
by: Huang, Yue, et al.
Published: (2026)
by: Huang, Yue, et al.
Published: (2026)
Planning to Explore: Curiosity-Driven Planning for LLM Test Generation
by: Amayuelas, Alfonso, et al.
Published: (2026)
by: Amayuelas, Alfonso, et al.
Published: (2026)
Similar Items
-
Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
by: Lin, Jiahang, et al.
Published: (2026) -
Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning
by: Liu, Yujian, et al.
Published: (2024) -
How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings
by: Liu, Yujian, et al.
Published: (2026) -
Absorber LLM: Harnessing Causal Synchronization for Test-Time Training
by: Zhang, Zhixin, et al.
Published: (2026) -
Revisiting Who's Harry Potter: Towards Targeted Unlearning from a Causal Intervention Perspective
by: Liu, Yujian, et al.
Published: (2024)