Guardado en:
| Autores principales: | Che, Sicong, Yang, Jiayi, Khurshid, Sarfraz, Wang, Wenxi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2603.00044 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
APRIL: API Synthesis with Automatic Prompt Optimization and Reinforcement Learning
por: Zhong, Hua, et al.
Publicado: (2025)
por: Zhong, Hua, et al.
Publicado: (2025)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
por: Xie, Zichen, et al.
Publicado: (2026)
por: Xie, Zichen, et al.
Publicado: (2026)
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
por: Xie, Zichen, et al.
Publicado: (2026)
por: Xie, Zichen, et al.
Publicado: (2026)
Generating executable oracles to check conformance of client code to requirements of JDK Javadocs using LLMs
por: Jiang, Shan, et al.
Publicado: (2024)
por: Jiang, Shan, et al.
Publicado: (2024)
OBsmith: LLM-Powered JavaScript Obfuscator Testing
por: Jiang, Shan, et al.
Publicado: (2025)
por: Jiang, Shan, et al.
Publicado: (2025)
MPBMC: Multi-Property Bounded Model Checking with GNN-guided Clustering
por: Roy, Soumik Guha, et al.
Publicado: (2026)
por: Roy, Soumik Guha, et al.
Publicado: (2026)
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
por: Zhao, Zhimin, et al.
Publicado: (2026)
por: Zhao, Zhimin, et al.
Publicado: (2026)
A Comprehensive Framework for Evaluating API-oriented Code Generation in Large Language Models
por: Wu, Yixi, et al.
Publicado: (2024)
por: Wu, Yixi, et al.
Publicado: (2024)
SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
por: Castellani, Tommaso, et al.
Publicado: (2025)
por: Castellani, Tommaso, et al.
Publicado: (2025)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
por: Bodla, Krishna Vamshi, et al.
Publicado: (2025)
por: Bodla, Krishna Vamshi, et al.
Publicado: (2025)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
por: Rodriguez-Cardenas, Daniel, et al.
Publicado: (2025)
por: Rodriguez-Cardenas, Daniel, et al.
Publicado: (2025)
A Large-Scale Study of Model Integration in ML-Enabled Software Systems
por: Sens, Yorick, et al.
Publicado: (2024)
por: Sens, Yorick, et al.
Publicado: (2024)
Code Generation by Differential Test Time Scaling
por: He, Yifeng, et al.
Publicado: (2026)
por: He, Yifeng, et al.
Publicado: (2026)
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
por: Bai, Yifan, et al.
Publicado: (2026)
por: Bai, Yifan, et al.
Publicado: (2026)
ManiTwin: Scaling Data-Generation-Ready Digital Object Dataset to 100K
por: Wang, Kaixuan, et al.
Publicado: (2026)
por: Wang, Kaixuan, et al.
Publicado: (2026)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
por: Daghighfarsoodeh, Alireza, et al.
Publicado: (2025)
por: Daghighfarsoodeh, Alireza, et al.
Publicado: (2025)
GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
por: Andrews, Martin, et al.
Publicado: (2025)
por: Andrews, Martin, et al.
Publicado: (2025)
On the Effectiveness of Large Language Models in Writing Alloy Formulas
por: Hong, Yang, et al.
Publicado: (2025)
por: Hong, Yang, et al.
Publicado: (2025)
EGSS: Entropy-guided Stepwise Scaling for Reliable Software Engineering
por: Mao, Chenhui, et al.
Publicado: (2026)
por: Mao, Chenhui, et al.
Publicado: (2026)
Understanding LLM-Driven Test Oracle Generation
por: Bodicoat, Adam, et al.
Publicado: (2026)
por: Bodicoat, Adam, et al.
Publicado: (2026)
A Theoretical Analysis of Test-Driven Code Generation
por: Menet, Nicolas, et al.
Publicado: (2026)
por: Menet, Nicolas, et al.
Publicado: (2026)
DeepKnowledge: Generalisation-Driven Deep Learning Testing
por: Missaoui, Sondess, et al.
Publicado: (2024)
por: Missaoui, Sondess, et al.
Publicado: (2024)
Towards a Classification of Open-Source ML Models and Datasets for Software Engineering
por: González, Alexandra, et al.
Publicado: (2024)
por: González, Alexandra, et al.
Publicado: (2024)
Scaling Down to Scale Up: A Cost-Benefit Analysis of Replacing OpenAI's LLM with Open Source SLMs in Production
por: Irugalbandara, Chandra, et al.
Publicado: (2023)
por: Irugalbandara, Chandra, et al.
Publicado: (2023)
miniCodeProps: a Minimal Benchmark for Proving Code Properties
por: Lohn, Evan, et al.
Publicado: (2024)
por: Lohn, Evan, et al.
Publicado: (2024)
A Model-Driven Engineering Approach to AI-Powered Healthcare Platforms
por: Raheem, Mira, et al.
Publicado: (2025)
por: Raheem, Mira, et al.
Publicado: (2025)
Protocol-Driven Development: Governing Generated Software Through Invariants and Continuous Evidence
por: He, Jun, et al.
Publicado: (2026)
por: He, Jun, et al.
Publicado: (2026)
Code Reborn AI-Driven Legacy Systems Modernization from COBOL to Java
por: Bandarupalli, Gopichand
Publicado: (2025)
por: Bandarupalli, Gopichand
Publicado: (2025)
Sketch-and-Verify: Structured Inference-Time Scaling via Program Sketching
por: Jiang, Shan, et al.
Publicado: (2026)
por: Jiang, Shan, et al.
Publicado: (2026)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
por: Ding, Yifeng, et al.
Publicado: (2026)
por: Ding, Yifeng, et al.
Publicado: (2026)
BitsAI-Fix: LLM-Driven Approach for Automated Lint Error Resolution in Practice
por: Li, Yuanpeng, et al.
Publicado: (2025)
por: Li, Yuanpeng, et al.
Publicado: (2025)
AI-Driven Code Refactoring: Using Graph Neural Networks to Enhance Software Maintainability
por: Bandarupalli, Gopichand
Publicado: (2025)
por: Bandarupalli, Gopichand
Publicado: (2025)
CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation
por: Wang, Xinchen, et al.
Publicado: (2025)
por: Wang, Xinchen, et al.
Publicado: (2025)
A Reference Architecture of Reinforcement Learning Frameworks
por: Liu, Xiaoran, et al.
Publicado: (2026)
por: Liu, Xiaoran, et al.
Publicado: (2026)
A Framework to Model ML Engineering Processes
por: Morales, Sergio, et al.
Publicado: (2024)
por: Morales, Sergio, et al.
Publicado: (2024)
Evaluating the Use of LLMs for Documentation to Code Traceability
por: Alor, Ebube, et al.
Publicado: (2025)
por: Alor, Ebube, et al.
Publicado: (2025)
Data-Driven Methods and AI in Engineering Design: A Systematic Literature Review Focusing on Challenges and Opportunities
por: Afifi, Nehal, et al.
Publicado: (2025)
por: Afifi, Nehal, et al.
Publicado: (2025)
Monitizer: Automating Design and Evaluation of Neural Network Monitors
por: Azeem, Muqsit, et al.
Publicado: (2024)
por: Azeem, Muqsit, et al.
Publicado: (2024)
OpenClassGen: A Large-Scale Corpus of Real-World Python Classes for LLM Research
por: Rahman, Musfiqur, et al.
Publicado: (2025)
por: Rahman, Musfiqur, et al.
Publicado: (2025)
Order Matters! An Empirical Study on Large Language Models' Input Order Bias in Software Fault Localization
por: Rafi, Md Nakhla, et al.
Publicado: (2024)
por: Rafi, Md Nakhla, et al.
Publicado: (2024)
Ejemplares similares
-
APRIL: API Synthesis with Automatic Prompt Optimization and Reinforcement Learning
por: Zhong, Hua, et al.
Publicado: (2025) -
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
por: Xie, Zichen, et al.
Publicado: (2026) -
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
por: Xie, Zichen, et al.
Publicado: (2026) -
Generating executable oracles to check conformance of client code to requirements of JDK Javadocs using LLMs
por: Jiang, Shan, et al.
Publicado: (2024) -
OBsmith: LLM-Powered JavaScript Obfuscator Testing
por: Jiang, Shan, et al.
Publicado: (2025)