Planning to Explore: Curiosity-Driven Planning for LLM Test Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Amayuelas, Alfonso, Laakom, Firas, Piękos, Piotr, Wang, Wenyi, Xu, Yifan, Wang, Yuhui, Schmidhuber, Jürgen, Wang, William |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors
par: Chen, Yimeng, et autres
Publié: (2025)
par: Chen, Yimeng, et autres
Publié: (2025)
CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging
par: Islam, Md. Ashraful, et autres
Publié: (2025)
par: Islam, Md. Ashraful, et autres
Publié: (2025)
The Prompt Alchemist: Automated LLM-Tailored Prompt Optimization for Test Case Generation
par: Gao, Shuzheng, et autres
Publié: (2025)
par: Gao, Shuzheng, et autres
Publié: (2025)
Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
par: Wang, Wenyi, et autres
Publié: (2025)
par: Wang, Wenyi, et autres
Publié: (2025)
PERC: Plan-As-Query Example Retrieval for Underrepresented Code Generation
par: Yoo, Jaeseok, et autres
Publié: (2024)
par: Yoo, Jaeseok, et autres
Publié: (2024)
Effective LLM-Driven Code Generation with Pythoness
par: Levin, Kyla H., et autres
Publié: (2025)
par: Levin, Kyla H., et autres
Publié: (2025)
RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
par: Luo, Jane, et autres
Publié: (2025)
par: Luo, Jane, et autres
Publié: (2025)
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
par: Zhang, Ziyao, et autres
Publié: (2024)
par: Zhang, Ziyao, et autres
Publié: (2024)
Reverse Chain: A Generic-Rule for LLMs to Master Multi-API Planning
par: Zhang, Yinger, et autres
Publié: (2023)
par: Zhang, Yinger, et autres
Publié: (2023)
Evaluating LLMs on Sequential API Call Through Automated Test Generation
par: Huang, Yuheng, et autres
Publié: (2025)
par: Huang, Yuheng, et autres
Publié: (2025)
A Review of Prominent Paradigms for LLM-Based Agents: Tool Use (Including RAG), Planning, and Feedback Learning
par: Li, Xinzhe
Publié: (2024)
par: Li, Xinzhe
Publié: (2024)
Curiosity-Driven Testing for Sequential Decision-Making Process
par: He, Junda, et autres
Publié: (2025)
par: He, Junda, et autres
Publié: (2025)
Text2Scenario: Text-Driven Scenario Generation for Autonomous Driving Test
par: Cai, Xuan, et autres
Publié: (2025)
par: Cai, Xuan, et autres
Publié: (2025)
From Code Generation to Software Testing: AI Copilot with Context-Based RAG
par: Wang, Yuchen, et autres
Publié: (2025)
par: Wang, Yuchen, et autres
Publié: (2025)
Evaluating Plan Compliance in Autonomous Programming Agents
par: Liu, Shuyang, et autres
Publié: (2026)
par: Liu, Shuyang, et autres
Publié: (2026)
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
par: Zhang, Haonan, et autres
Publié: (2025)
par: Zhang, Haonan, et autres
Publié: (2025)
Testing and Evaluation of Large Language Models: Correctness, Non-Toxicity, and Fairness
par: Wang, Wenxuan
Publié: (2024)
par: Wang, Wenxuan
Publié: (2024)
FACTS: A Factored State-Space Framework For World Modelling
par: Nanbo, Li, et autres
Publié: (2024)
par: Nanbo, Li, et autres
Publié: (2024)
AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection
par: Zhao, Zijie, et autres
Publié: (2026)
par: Zhao, Zijie, et autres
Publié: (2026)
Validating LLM-Generated Programs with Metamorphic Prompt Testing
par: Wang, Xiaoyin, et autres
Publié: (2024)
par: Wang, Xiaoyin, et autres
Publié: (2024)
OBsmith: LLM-Powered JavaScript Obfuscator Testing
par: Jiang, Shan, et autres
Publié: (2025)
par: Jiang, Shan, et autres
Publié: (2025)
Illocutionary Explanation Planning for Source-Faithful Explanations in Retrieval-Augmented Language Models
par: Sovrano, Francesco, et autres
Publié: (2026)
par: Sovrano, Francesco, et autres
Publié: (2026)
SPARC: Scenario Planning and Reasoning for Automated C Unit Test Generation
par: Chowdhury, Jaid Monwar, et autres
Publié: (2026)
par: Chowdhury, Jaid Monwar, et autres
Publié: (2026)
ReFuzzer: Feedback-Driven Approach to Enhance Validity of LLM-Generated Test Programs
par: Shree, Iti, et autres
Publié: (2025)
par: Shree, Iti, et autres
Publié: (2025)
Learner-Tailored Program Repair: A Solution Generator with Iterative Edit-Driven Retrieval Enhancement
par: Dai, Zhenlong, et autres
Publié: (2026)
par: Dai, Zhenlong, et autres
Publié: (2026)
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
par: Cui, Yi
Publié: (2025)
par: Cui, Yi
Publié: (2025)
MutaGReP: Execution-Free Repository-Grounded Plan Search for Code-Use
par: Khan, Zaid, et autres
Publié: (2025)
par: Khan, Zaid, et autres
Publié: (2025)
Call-Chain-Aware LLM-Based Test Generation for Java Projects
par: Wang, Guancheng, et autres
Publié: (2026)
par: Wang, Guancheng, et autres
Publié: (2026)
Agentic Interpretation: Lattice-Structured Evidence for LLM-Based Program Analysis
par: Mitchell, Jacqueline L., et autres
Publié: (2026)
par: Mitchell, Jacqueline L., et autres
Publié: (2026)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
par: Zhang, William, et autres
Publié: (2024)
par: Zhang, William, et autres
Publié: (2024)
Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems
par: Cao, Yuhan, et autres
Publié: (2025)
par: Cao, Yuhan, et autres
Publié: (2025)
Issue Localization via LLM-Driven Iterative Code Graph Searching
par: Jiang, Zhonghao, et autres
Publié: (2025)
par: Jiang, Zhonghao, et autres
Publié: (2025)
Polygon: Symbolic Reasoning for SQL using Conflict-Driven Under-Approximation Search
par: Zhao, Pinhan, et autres
Publié: (2025)
par: Zhao, Pinhan, et autres
Publié: (2025)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
par: Lai, Peng, et autres
Publié: (2026)
par: Lai, Peng, et autres
Publié: (2026)
Testing the Effect of Code Documentation on Large Language Model Code Understanding
par: Macke, William, et autres
Publié: (2024)
par: Macke, William, et autres
Publié: (2024)
Crystal: Illuminating LLM Abilities on Language and Code
par: Tao, Tianhua, et autres
Publié: (2024)
par: Tao, Tianhua, et autres
Publié: (2024)
IndustryCode: A Benchmark for Industry Code Generation
par: Zeng, Puyu, et autres
Publié: (2026)
par: Zeng, Puyu, et autres
Publié: (2026)
AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT Applications
par: Shen, Leming, et autres
Publié: (2025)
par: Shen, Leming, et autres
Publié: (2025)
Dynamic Stability of LLM-Generated Code
par: Rajput, Prateek, et autres
Publié: (2025)
par: Rajput, Prateek, et autres
Publié: (2025)
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
par: Wang, Yibo, et autres
Publié: (2025)
par: Wang, Yibo, et autres
Publié: (2025)
Documents similaires
-
PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors
par: Chen, Yimeng, et autres
Publié: (2025) -
CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging
par: Islam, Md. Ashraful, et autres
Publié: (2025) -
The Prompt Alchemist: Automated LLM-Tailored Prompt Optimization for Test Case Generation
par: Gao, Shuzheng, et autres
Publié: (2025) -
Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
par: Wang, Wenyi, et autres
Publié: (2025) -
PERC: Plan-As-Query Example Retrieval for Underrepresented Code Generation
par: Yoo, Jaeseok, et autres
Publié: (2024)