Heterogeneous Prompting and Execution Feedback for SWE Issue Test Generation and Selection
Fuente:
arXiv
Saved in:
| Main Authors: | Ahmed, Toufique, Ganhotra, Jatin, Shinnar, Avraham, Hirzel, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reproduction Test Generation for Java SWE Issues
by: Ahmed, Toufique, et al.
Published: (2026)
by: Ahmed, Toufique, et al.
Published: (2026)
Investigating Test Overfitting on SWE-bench
by: Ahmed, Toufique, et al.
Published: (2025)
by: Ahmed, Toufique, et al.
Published: (2025)
Otter: Generating Tests from Issues to Validate SWE Patches
by: Ahmed, Toufique, et al.
Published: (2025)
by: Ahmed, Toufique, et al.
Published: (2025)
Resolving Java Code Repository Issues with iSWE Agent
by: Ganhotra, Jatin, et al.
Published: (2026)
by: Ganhotra, Jatin, et al.
Published: (2026)
TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
by: Ahmed, Toufique, et al.
Published: (2024)
by: Ahmed, Toufique, et al.
Published: (2024)
Can Old Tests Do New Tricks for Resolving SWE Issues?
by: Chen, Yang, et al.
Published: (2025)
by: Chen, Yang, et al.
Published: (2025)
Evaluating Plan Compliance in Autonomous Programming Agents
by: Liu, Shuyang, et al.
Published: (2026)
by: Liu, Shuyang, et al.
Published: (2026)
When Agents go Astray: Course-Correcting SWE Agents with PRMs
by: Gandhi, Shubham, et al.
Published: (2025)
by: Gandhi, Shubham, et al.
Published: (2025)
Echo: Graph-Enhanced Retrieval and Execution Feedback for Issue Reproduction Test Generation
by: Fei, Zhiwei, et al.
Published: (2026)
by: Fei, Zhiwei, et al.
Published: (2026)
From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents
by: Ludwig, Nikolai, et al.
Published: (2026)
by: Ludwig, Nikolai, et al.
Published: (2026)
Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization)
by: Ahmed, Toufique, et al.
Published: (2023)
by: Ahmed, Toufique, et al.
Published: (2023)
TestWeaver: Execution-aware, Feedback-driven Regression Testing Generation with Large Language Models
by: Le, Cuong Chi, et al.
Published: (2025)
by: Le, Cuong Chi, et al.
Published: (2025)
GenX: Mastering Code and Test Generation with Execution Feedback
by: Wang, Nan, et al.
Published: (2024)
by: Wang, Nan, et al.
Published: (2024)
CoDocBench: A Dataset for Code-Documentation Alignment in Software Maintenance
by: Pai, Kunal, et al.
Published: (2025)
by: Pai, Kunal, et al.
Published: (2025)
Calibration of Large Language Models on Code Summarization
by: Virk, Yuvraj, et al.
Published: (2024)
by: Virk, Yuvraj, et al.
Published: (2024)
Process-Centric Analysis of Agentic Software Systems
by: Liu, Shuyang, et al.
Published: (2025)
by: Liu, Shuyang, et al.
Published: (2025)
SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation
by: Oliva, Gustavo A., et al.
Published: (2025)
by: Oliva, Gustavo A., et al.
Published: (2025)
SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling
by: Han, Hao, et al.
Published: (2026)
by: Han, Hao, et al.
Published: (2026)
SWE-Manager: Selecting and Synthesizing Golden Proposals Before Coding
by: Tan, Boyin, et al.
Published: (2026)
by: Tan, Boyin, et al.
Published: (2026)
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
by: Prathifkumar, Thanosan, et al.
Published: (2025)
by: Prathifkumar, Thanosan, et al.
Published: (2025)
SWE-Mirror: Scaling Issue-Resolving Datasets by Mirroring Issues Across Repositories
by: Wang, Junhao, et al.
Published: (2025)
by: Wang, Junhao, et al.
Published: (2025)
Prompting and Fine-tuning Large Language Models for Automated Code Review Comment Generation
by: Haider, Md. Asif, et al.
Published: (2024)
by: Haider, Md. Asif, et al.
Published: (2024)
Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study
by: Wang, You, et al.
Published: (2025)
by: Wang, You, et al.
Published: (2025)
Studying LLM Performance on Closed- and Open-source Data
by: Ahmed, Toufique, et al.
Published: (2024)
by: Ahmed, Toufique, et al.
Published: (2024)
SWE-Hub: A Unified Production System for Scalable, Executable Software Engineering Tasks
by: Zeng, Yucheng, et al.
Published: (2026)
by: Zeng, Yucheng, et al.
Published: (2026)
Mokav: Execution-driven Differential Testing with LLMs
by: Etemadi, Khashayar, et al.
Published: (2024)
by: Etemadi, Khashayar, et al.
Published: (2024)
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle
by: Guan, Hao, et al.
Published: (2026)
by: Guan, Hao, et al.
Published: (2026)
Sifting through the Chaff: On Utilizing Execution Feedback for Ranking the Generated Code Candidates
by: Sun, Zhihong, et al.
Published: (2024)
by: Sun, Zhihong, et al.
Published: (2024)
Improving Examples in Web API Specifications using Iterated-Calls In-Context Learning
by: Jain, Kush, et al.
Published: (2025)
by: Jain, Kush, et al.
Published: (2025)
TestForge: Feedback-Driven, Agentic Test Suite Generation
by: Jain, Kush, et al.
Published: (2025)
by: Jain, Kush, et al.
Published: (2025)
SWE-PRBench: Benchmarking AI Code Review Quality Against Pull Request Feedback
by: Kumar, Deepak
Published: (2026)
by: Kumar, Deepak
Published: (2026)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
by: Aleithan, Reem, et al.
Published: (2024)
by: Aleithan, Reem, et al.
Published: (2024)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
by: Raghavendra, Mohit, et al.
Published: (2026)
by: Raghavendra, Mohit, et al.
Published: (2026)
SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale
by: Badertdinov, Ibragim, et al.
Published: (2026)
by: Badertdinov, Ibragim, et al.
Published: (2026)
SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents
by: Dihan, Mahir Labib, et al.
Published: (2026)
by: Dihan, Mahir Labib, et al.
Published: (2026)
Scaling Mobile Chaos Testing with AI-Driven Test Execution
by: Marcano, Juan, et al.
Published: (2026)
by: Marcano, Juan, et al.
Published: (2026)
Test Wars: A Comparative Study of SBST, Symbolic Execution, and LLM-Based Approaches to Unit Test Generation
by: Abdullin, Azat, et al.
Published: (2025)
by: Abdullin, Azat, et al.
Published: (2025)
PPO guided Agentic Pipeline for Adaptive Prompt Selection and Test Case Generation
by: Koushik, Gourisetty Venkata Sai, et al.
Published: (2026)
by: Koushik, Gourisetty Venkata Sai, et al.
Published: (2026)
GHIssuemarket: A Sandbox Environment for SWE-Agents Economic Experimentation
by: Fouad, Mohamed A., et al.
Published: (2024)
by: Fouad, Mohamed A., et al.
Published: (2024)
Representing Prompting Patterns with PDL: Compliance Agent Case Study
by: Vaziri, Mandana, et al.
Published: (2025)
by: Vaziri, Mandana, et al.
Published: (2025)
Similar Items
-
Reproduction Test Generation for Java SWE Issues
by: Ahmed, Toufique, et al.
Published: (2026) -
Investigating Test Overfitting on SWE-bench
by: Ahmed, Toufique, et al.
Published: (2025) -
Otter: Generating Tests from Issues to Validate SWE Patches
by: Ahmed, Toufique, et al.
Published: (2025) -
Resolving Java Code Repository Issues with iSWE Agent
by: Ganhotra, Jatin, et al.
Published: (2026) -
TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
by: Ahmed, Toufique, et al.
Published: (2024)