Artisan: Agentic Artifact Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Baek, Doehyun, Pradel, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Execution-Aware Program Reduction for WebAssembly via Record and Replay
von: Baek, Doehyun, et al.
Veröffentlicht: (2025)
von: Baek, Doehyun, et al.
Veröffentlicht: (2025)
Testora: Using Natural Language Intent to Detect Behavioral Regressions
von: Pradel, Michael
Veröffentlicht: (2025)
von: Pradel, Michael
Veröffentlicht: (2025)
PatchGuru: Patch Oracle Inference from Natural Language Artifacts with Large Language Models
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2026)
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2026)
Evaluating LLM Agents on Automated Software Analysis Tasks
von: Bouzenia, Islem, et al.
Veröffentlicht: (2026)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2026)
Agentic AI Software Engineers: Programming with Trust
von: Roychoudhury, Abhik, et al.
Veröffentlicht: (2025)
von: Roychoudhury, Abhik, et al.
Veröffentlicht: (2025)
De-Hallucinator: Mitigating LLM Hallucinations in Code Generation Tasks via Iterative Grounding
von: Eghbali, Aryaz, et al.
Veröffentlicht: (2024)
von: Eghbali, Aryaz, et al.
Veröffentlicht: (2024)
CodeMapper: A Language-Agnostic Approach to Mapping Code Regions Across Commits
von: Hu, Huimin, et al.
Veröffentlicht: (2025)
von: Hu, Huimin, et al.
Veröffentlicht: (2025)
Can LLMs Replace Manual Annotation of Software Engineering Artifacts?
von: Ahmed, Toufique, et al.
Veröffentlicht: (2024)
von: Ahmed, Toufique, et al.
Veröffentlicht: (2024)
RippleGUItester: Change-Aware Exploratory Testing
von: Su, Yanqi, et al.
Veröffentlicht: (2026)
von: Su, Yanqi, et al.
Veröffentlicht: (2026)
Names Are All You Need: Effective and Safe Regression Test Selection for Python
von: Wang, You, et al.
Veröffentlicht: (2026)
von: Wang, You, et al.
Veröffentlicht: (2026)
Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
von: Wang, You, et al.
Veröffentlicht: (2025)
von: Wang, You, et al.
Veröffentlicht: (2025)
ChangeGuard: Validating Code Changes via Pairwise Learning-Guided Execution
von: Gröninger, Lars, et al.
Veröffentlicht: (2024)
von: Gröninger, Lars, et al.
Veröffentlicht: (2024)
NoCode-bench: A Benchmark for Evaluating Natural Language-Driven Feature Addition
von: Deng, Le, et al.
Veröffentlicht: (2025)
von: Deng, Le, et al.
Veröffentlicht: (2025)
AgentStepper: Interactive Debugging of Software Development Agents
von: Hutter, Robert, et al.
Veröffentlicht: (2026)
von: Hutter, Robert, et al.
Veröffentlicht: (2026)
Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories
von: Bouzenia, Islem, et al.
Veröffentlicht: (2025)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2025)
Treefix: Enabling Execution with a Tree of Prefixes
von: Souza, Beatriz, et al.
Veröffentlicht: (2025)
von: Souza, Beatriz, et al.
Veröffentlicht: (2025)
You Name It, I Run It: An LLM Agent to Execute Tests of Arbitrary Projects
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
DyPyBench: A Benchmark of Executable Python Software
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
Analyzing Quantum Programs with LintQ: A Static Analysis Framework for Qiskit
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2023)
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2023)
Agent-Based Software Artifact Evaluation
von: Wu, Zhaonan, et al.
Veröffentlicht: (2026)
von: Wu, Zhaonan, et al.
Veröffentlicht: (2026)
FlyCatcher: Neural Inference of Runtime Checkers from Tests
von: Souza, Beatriz, et al.
Veröffentlicht: (2026)
von: Souza, Beatriz, et al.
Veröffentlicht: (2026)
Change And Cover: Last-Mile, Pull Request-Based Regression Test Augmentation
von: Zhou, Zitong, et al.
Veröffentlicht: (2026)
von: Zhou, Zitong, et al.
Veröffentlicht: (2026)
Issue2Test: Generating Reproducing Test Cases from Issue Reports
von: Nashid, Noor, et al.
Veröffentlicht: (2025)
von: Nashid, Noor, et al.
Veröffentlicht: (2025)
PyTy: Repairing Static Type Errors in Python
von: Chow, Yiu Wai, et al.
Veröffentlicht: (2024)
von: Chow, Yiu Wai, et al.
Veröffentlicht: (2024)
CodeCureAgent: Automatic Classification and Repair of Static Analysis Warnings
von: Joos, Pascal, et al.
Veröffentlicht: (2025)
von: Joos, Pascal, et al.
Veröffentlicht: (2025)
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
von: Bouzenia, Islem, et al.
Veröffentlicht: (2024)
PoCGen: Generating Proof-of-Concept Exploits for Vulnerabilities in Npm Packages
von: Simsek, Deniz, et al.
Veröffentlicht: (2025)
von: Simsek, Deniz, et al.
Veröffentlicht: (2025)
Rethinking Artifact Evaluation for Software Engineering in the Age of Generative AI
von: Treude, Christoph, et al.
Veröffentlicht: (2026)
von: Treude, Christoph, et al.
Veröffentlicht: (2026)
From Bugs to Benchmarks: A Comprehensive Survey of Software Defect Datasets
von: Zhu, Hao-Nan, et al.
Veröffentlicht: (2025)
von: Zhu, Hao-Nan, et al.
Veröffentlicht: (2025)
QITE: Assembly-Level, Cross-Platform Testing of Quantum Computing Platforms
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2025)
von: Paltenghi, Matteo, et al.
Veröffentlicht: (2025)
Auxiliary Artifacts in Requirements Traceability: A Systematic Mapping Study
von: Abdeen, Waleed, et al.
Veröffentlicht: (2025)
von: Abdeen, Waleed, et al.
Veröffentlicht: (2025)
On Developing an Artifact-based Approach to Regulatory Requirements Engineering
von: Kosenkov, Oleksandr, et al.
Veröffentlicht: (2024)
von: Kosenkov, Oleksandr, et al.
Veröffentlicht: (2024)
Requirements Quality Research Artifacts: Recovery, Analysis, and Management Guideline
von: Frattini, Julian, et al.
Veröffentlicht: (2024)
von: Frattini, Julian, et al.
Veröffentlicht: (2024)
An Agentic Approach Towards Replication Package Quality Evaluation
von: Mbida, Maximilian Alexander Amougou, et al.
Veröffentlicht: (2026)
von: Mbida, Maximilian Alexander Amougou, et al.
Veröffentlicht: (2026)
Agentic Repository Mining: A Multi-Task Evaluation
von: Härtel, Johannes
Veröffentlicht: (2026)
von: Härtel, Johannes
Veröffentlicht: (2026)
LLM-Based Repair of Static Nullability Errors
von: Karimipour, Nima, et al.
Veröffentlicht: (2025)
von: Karimipour, Nima, et al.
Veröffentlicht: (2025)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
von: Fandina, Ora Nova, et al.
Veröffentlicht: (2025)
von: Fandina, Ora Nova, et al.
Veröffentlicht: (2025)
Intent Preserving Generation of Diverse and Idiomatic (Code-)Artifacts
von: Westphal, Oliver
Veröffentlicht: (2025)
von: Westphal, Oliver
Veröffentlicht: (2025)
Research Artifacts in Software Engineering Publications: Status and Trends
von: Liu, Mugeng, et al.
Veröffentlicht: (2024)
von: Liu, Mugeng, et al.
Veröffentlicht: (2024)
Fuzz4All: Universal Fuzzing with Large Language Models
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2023)
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Execution-Aware Program Reduction for WebAssembly via Record and Replay
von: Baek, Doehyun, et al.
Veröffentlicht: (2025) -
Testora: Using Natural Language Intent to Detect Behavioral Regressions
von: Pradel, Michael
Veröffentlicht: (2025) -
PatchGuru: Patch Oracle Inference from Natural Language Artifacts with Large Language Models
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2026) -
Evaluating LLM Agents on Automated Software Analysis Tasks
von: Bouzenia, Islem, et al.
Veröffentlicht: (2026) -
Agentic AI Software Engineers: Programming with Trust
von: Roychoudhury, Abhik, et al.
Veröffentlicht: (2025)