From Procedural Skills to Strategy Genes: Towards Experience-Driven Test-Time Evolution
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Junjie, Ren, Yiming, Zhang, Haoyang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EffiSkill: Agent Skill Based Automated Code Efficiency Optimization
di: Wang, Zimu, et al.
Pubblicazione: (2026)
di: Wang, Zimu, et al.
Pubblicazione: (2026)
Seeker: Towards Exception Safety Code Generation with Intermediate Language Agents Framework
di: Zhang, Xuanming, et al.
Pubblicazione: (2024)
di: Zhang, Xuanming, et al.
Pubblicazione: (2024)
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
di: Chen, Shiqi, et al.
Pubblicazione: (2026)
di: Chen, Shiqi, et al.
Pubblicazione: (2026)
Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
di: Lin, Jiahang, et al.
Pubblicazione: (2026)
di: Lin, Jiahang, et al.
Pubblicazione: (2026)
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
di: Xu, Haoyuan, et al.
Pubblicazione: (2026)
di: Xu, Haoyuan, et al.
Pubblicazione: (2026)
A Comparative Study on the Impact of Test-Driven Development (TDD) and Behavior-Driven Development (BDD) on Enterprise Software Delivery Effectiveness
di: Cui, Jun
Pubblicazione: (2024)
di: Cui, Jun
Pubblicazione: (2024)
Text2Scenario: Text-Driven Scenario Generation for Autonomous Driving Test
di: Cai, Xuan, et al.
Pubblicazione: (2025)
di: Cai, Xuan, et al.
Pubblicazione: (2025)
SWE-Exp: Experience-Driven Software Issue Resolution
di: Chen, Silin, et al.
Pubblicazione: (2025)
di: Chen, Silin, et al.
Pubblicazione: (2025)
Planning to Explore: Curiosity-Driven Planning for LLM Test Generation
di: Amayuelas, Alfonso, et al.
Pubblicazione: (2026)
di: Amayuelas, Alfonso, et al.
Pubblicazione: (2026)
Machine Translation Testing via Syntactic Tree Pruning
di: Zhang, Quanjun, et al.
Pubblicazione: (2024)
di: Zhang, Quanjun, et al.
Pubblicazione: (2024)
SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation
di: Ma, Zeyao, et al.
Pubblicazione: (2024)
di: Ma, Zeyao, et al.
Pubblicazione: (2024)
GUI Test Migration via Abstraction and Concretization
di: Zhang, Yakun, et al.
Pubblicazione: (2024)
di: Zhang, Yakun, et al.
Pubblicazione: (2024)
Dynamic Scaling of Unit Tests for Code Reward Modeling
di: Ma, Zeyao, et al.
Pubblicazione: (2025)
di: Ma, Zeyao, et al.
Pubblicazione: (2025)
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
di: Liu, Steven, et al.
Pubblicazione: (2026)
di: Liu, Steven, et al.
Pubblicazione: (2026)
Skill over Scale: The Case for Medium, Domain-Specific Models for SE
di: Mukherjee, Manisha, et al.
Pubblicazione: (2023)
di: Mukherjee, Manisha, et al.
Pubblicazione: (2023)
Solver-Independent Automated Problem Formulation via LLMs for High-Cost Simulation-Driven Design
di: Li, Yuchen, et al.
Pubblicazione: (2025)
di: Li, Yuchen, et al.
Pubblicazione: (2025)
Measuring the Influence of Incorrect Code on Test Generation
di: Huang, Dong, et al.
Pubblicazione: (2024)
di: Huang, Dong, et al.
Pubblicazione: (2024)
Sphinx: Benchmarking and Modeling for LLM-Driven Pull Request Review
di: Zhang, Daoan, et al.
Pubblicazione: (2026)
di: Zhang, Daoan, et al.
Pubblicazione: (2026)
Assessing Evaluation Metrics for Neural Test Oracle Generation
di: Shin, Jiho, et al.
Pubblicazione: (2023)
di: Shin, Jiho, et al.
Pubblicazione: (2023)
Nexus: Execution-Grounded Multi-Agent Test Oracle Synthesis
di: Huang, Dong, et al.
Pubblicazione: (2025)
di: Huang, Dong, et al.
Pubblicazione: (2025)
Leveraging LLMs for Grammar Adaptation: A Study on Metamodel-Grammar Co-Evolution
di: Zhang, Weixing, et al.
Pubblicazione: (2026)
di: Zhang, Weixing, et al.
Pubblicazione: (2026)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
di: Chen, Zaoyu, et al.
Pubblicazione: (2026)
di: Chen, Zaoyu, et al.
Pubblicazione: (2026)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
di: Huang, Dong, et al.
Pubblicazione: (2025)
di: Huang, Dong, et al.
Pubblicazione: (2025)
GenX: Mastering Code and Test Generation with Execution Feedback
di: Wang, Nan, et al.
Pubblicazione: (2024)
di: Wang, Nan, et al.
Pubblicazione: (2024)
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
di: Wang, Yibo, et al.
Pubblicazione: (2025)
di: Wang, Yibo, et al.
Pubblicazione: (2025)
Log Summarisation for Defect Evolution Analysis
di: Dolga, Rares, et al.
Pubblicazione: (2024)
di: Dolga, Rares, et al.
Pubblicazione: (2024)
MOSS: Enabling Code-Driven Evolution and Context Management for AI Agents
di: Zhu, Ming, et al.
Pubblicazione: (2024)
di: Zhu, Ming, et al.
Pubblicazione: (2024)
Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling
di: Jiao, Yizhu, et al.
Pubblicazione: (2026)
di: Jiao, Yizhu, et al.
Pubblicazione: (2026)
CodeContests+: High-Quality Test Case Generation for Competitive Programming
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
HarnessLLM: Automatic Testing Harness Generation via Reinforcement Learning
di: Liu, Yujian, et al.
Pubblicazione: (2025)
di: Liu, Yujian, et al.
Pubblicazione: (2025)
FireBench: Evaluating Instruction Following in Enterprise and API-Driven LLM Applications
di: Zhang, Yunfan, et al.
Pubblicazione: (2026)
di: Zhang, Yunfan, et al.
Pubblicazione: (2026)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
di: Fu, Lingyue, et al.
Pubblicazione: (2025)
DiffuTester: Accelerating Unit Test Generation for Diffusion LLMs via Mining Structural Pattern
di: Yang, Lekang, et al.
Pubblicazione: (2025)
di: Yang, Lekang, et al.
Pubblicazione: (2025)
From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning
di: Zhang, Jiajun, et al.
Pubblicazione: (2026)
di: Zhang, Jiajun, et al.
Pubblicazione: (2026)
SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development
di: Du, Yaxin, et al.
Pubblicazione: (2025)
di: Du, Yaxin, et al.
Pubblicazione: (2025)
Towards Exception Safety Code Generation with Intermediate Representation Agents Framework
di: Zhang, Xuanming, et al.
Pubblicazione: (2024)
di: Zhang, Xuanming, et al.
Pubblicazione: (2024)
Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability
di: Chen, Zejian, et al.
Pubblicazione: (2026)
di: Chen, Zejian, et al.
Pubblicazione: (2026)
Interpretable Online Log Analysis Using Large Language Models with Prompt Strategies
di: Liu, Yilun, et al.
Pubblicazione: (2023)
di: Liu, Yilun, et al.
Pubblicazione: (2023)
From Code Foundation Models to Agents and Applications: A Comprehensive Survey and Practical Guide to Code Intelligence
di: Yang, Jian, et al.
Pubblicazione: (2025)
di: Yang, Jian, et al.
Pubblicazione: (2025)
VersiCode: Towards Version-controllable Code Generation
di: Wu, Tongtong, et al.
Pubblicazione: (2024)
di: Wu, Tongtong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
EffiSkill: Agent Skill Based Automated Code Efficiency Optimization
di: Wang, Zimu, et al.
Pubblicazione: (2026) -
Seeker: Towards Exception Safety Code Generation with Intermediate Language Agents Framework
di: Zhang, Xuanming, et al.
Pubblicazione: (2024) -
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
di: Chen, Shiqi, et al.
Pubblicazione: (2026) -
Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
di: Lin, Jiahang, et al.
Pubblicazione: (2026) -
The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration
di: Xu, Haoyuan, et al.
Pubblicazione: (2026)