Saved in:
| Main Authors: | Che, Sicong, Yang, Jiayi, Khurshid, Sarfraz, Wang, Wenxi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.00044 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
APRIL: API Synthesis with Automatic Prompt Optimization and Reinforcement Learning
by: Zhong, Hua, et al.
Published: (2025)
by: Zhong, Hua, et al.
Published: (2025)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
by: Xie, Zichen, et al.
Published: (2026)
by: Xie, Zichen, et al.
Published: (2026)
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
by: Xie, Zichen, et al.
Published: (2026)
by: Xie, Zichen, et al.
Published: (2026)
Generating executable oracles to check conformance of client code to requirements of JDK Javadocs using LLMs
by: Jiang, Shan, et al.
Published: (2024)
by: Jiang, Shan, et al.
Published: (2024)
OBsmith: LLM-Powered JavaScript Obfuscator Testing
by: Jiang, Shan, et al.
Published: (2025)
by: Jiang, Shan, et al.
Published: (2025)
MPBMC: Multi-Property Bounded Model Checking with GNN-guided Clustering
by: Roy, Soumik Guha, et al.
Published: (2026)
by: Roy, Soumik Guha, et al.
Published: (2026)
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
by: Zhao, Zhimin, et al.
Published: (2026)
by: Zhao, Zhimin, et al.
Published: (2026)
A Comprehensive Framework for Evaluating API-oriented Code Generation in Large Language Models
by: Wu, Yixi, et al.
Published: (2024)
by: Wu, Yixi, et al.
Published: (2024)
SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
by: Castellani, Tommaso, et al.
Published: (2025)
by: Castellani, Tommaso, et al.
Published: (2025)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
by: Bodla, Krishna Vamshi, et al.
Published: (2025)
by: Bodla, Krishna Vamshi, et al.
Published: (2025)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2025)
by: Rodriguez-Cardenas, Daniel, et al.
Published: (2025)
A Large-Scale Study of Model Integration in ML-Enabled Software Systems
by: Sens, Yorick, et al.
Published: (2024)
by: Sens, Yorick, et al.
Published: (2024)
Code Generation by Differential Test Time Scaling
by: He, Yifeng, et al.
Published: (2026)
by: He, Yifeng, et al.
Published: (2026)
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
by: Bai, Yifan, et al.
Published: (2026)
by: Bai, Yifan, et al.
Published: (2026)
ManiTwin: Scaling Data-Generation-Ready Digital Object Dataset to 100K
by: Wang, Kaixuan, et al.
Published: (2026)
by: Wang, Kaixuan, et al.
Published: (2026)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
by: Andrews, Martin, et al.
Published: (2025)
by: Andrews, Martin, et al.
Published: (2025)
On the Effectiveness of Large Language Models in Writing Alloy Formulas
by: Hong, Yang, et al.
Published: (2025)
by: Hong, Yang, et al.
Published: (2025)
EGSS: Entropy-guided Stepwise Scaling for Reliable Software Engineering
by: Mao, Chenhui, et al.
Published: (2026)
by: Mao, Chenhui, et al.
Published: (2026)
Understanding LLM-Driven Test Oracle Generation
by: Bodicoat, Adam, et al.
Published: (2026)
by: Bodicoat, Adam, et al.
Published: (2026)
A Theoretical Analysis of Test-Driven Code Generation
by: Menet, Nicolas, et al.
Published: (2026)
by: Menet, Nicolas, et al.
Published: (2026)
DeepKnowledge: Generalisation-Driven Deep Learning Testing
by: Missaoui, Sondess, et al.
Published: (2024)
by: Missaoui, Sondess, et al.
Published: (2024)
Towards a Classification of Open-Source ML Models and Datasets for Software Engineering
by: González, Alexandra, et al.
Published: (2024)
by: González, Alexandra, et al.
Published: (2024)
Scaling Down to Scale Up: A Cost-Benefit Analysis of Replacing OpenAI's LLM with Open Source SLMs in Production
by: Irugalbandara, Chandra, et al.
Published: (2023)
by: Irugalbandara, Chandra, et al.
Published: (2023)
miniCodeProps: a Minimal Benchmark for Proving Code Properties
by: Lohn, Evan, et al.
Published: (2024)
by: Lohn, Evan, et al.
Published: (2024)
A Model-Driven Engineering Approach to AI-Powered Healthcare Platforms
by: Raheem, Mira, et al.
Published: (2025)
by: Raheem, Mira, et al.
Published: (2025)
Protocol-Driven Development: Governing Generated Software Through Invariants and Continuous Evidence
by: He, Jun, et al.
Published: (2026)
by: He, Jun, et al.
Published: (2026)
Code Reborn AI-Driven Legacy Systems Modernization from COBOL to Java
by: Bandarupalli, Gopichand
Published: (2025)
by: Bandarupalli, Gopichand
Published: (2025)
Sketch-and-Verify: Structured Inference-Time Scaling via Program Sketching
by: Jiang, Shan, et al.
Published: (2026)
by: Jiang, Shan, et al.
Published: (2026)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
by: Ding, Yifeng, et al.
Published: (2026)
by: Ding, Yifeng, et al.
Published: (2026)
BitsAI-Fix: LLM-Driven Approach for Automated Lint Error Resolution in Practice
by: Li, Yuanpeng, et al.
Published: (2025)
by: Li, Yuanpeng, et al.
Published: (2025)
AI-Driven Code Refactoring: Using Graph Neural Networks to Enhance Software Maintainability
by: Bandarupalli, Gopichand
Published: (2025)
by: Bandarupalli, Gopichand
Published: (2025)
CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation
by: Wang, Xinchen, et al.
Published: (2025)
by: Wang, Xinchen, et al.
Published: (2025)
A Reference Architecture of Reinforcement Learning Frameworks
by: Liu, Xiaoran, et al.
Published: (2026)
by: Liu, Xiaoran, et al.
Published: (2026)
A Framework to Model ML Engineering Processes
by: Morales, Sergio, et al.
Published: (2024)
by: Morales, Sergio, et al.
Published: (2024)
Evaluating the Use of LLMs for Documentation to Code Traceability
by: Alor, Ebube, et al.
Published: (2025)
by: Alor, Ebube, et al.
Published: (2025)
Data-Driven Methods and AI in Engineering Design: A Systematic Literature Review Focusing on Challenges and Opportunities
by: Afifi, Nehal, et al.
Published: (2025)
by: Afifi, Nehal, et al.
Published: (2025)
Monitizer: Automating Design and Evaluation of Neural Network Monitors
by: Azeem, Muqsit, et al.
Published: (2024)
by: Azeem, Muqsit, et al.
Published: (2024)
OpenClassGen: A Large-Scale Corpus of Real-World Python Classes for LLM Research
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
Order Matters! An Empirical Study on Large Language Models' Input Order Bias in Software Fault Localization
by: Rafi, Md Nakhla, et al.
Published: (2024)
by: Rafi, Md Nakhla, et al.
Published: (2024)
Similar Items
-
APRIL: API Synthesis with Automatic Prompt Optimization and Reinforcement Learning
by: Zhong, Hua, et al.
Published: (2025) -
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
by: Xie, Zichen, et al.
Published: (2026) -
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
by: Xie, Zichen, et al.
Published: (2026) -
Generating executable oracles to check conformance of client code to requirements of JDK Javadocs using LLMs
by: Jiang, Shan, et al.
Published: (2024) -
OBsmith: LLM-Powered JavaScript Obfuscator Testing
by: Jiang, Shan, et al.
Published: (2025)