SWE-Shepherd: Advancing PRMs for Reinforcing Code Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Dihan, Mahir Labib, Khan, Md Ashrafur Rahman |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BanglaForge: LLM Collaboration with Self-Refinement for Bangla Code Generation
by: Dihan, Mahir Labib, et al.
Published: (2025)
by: Dihan, Mahir Labib, et al.
Published: (2025)
LLMSniffer: Detecting LLM-Generated Code via GraphCodeBERT and Supervised Contrastive Learning
by: Dihan, Mahir Labib, et al.
Published: (2026)
by: Dihan, Mahir Labib, et al.
Published: (2026)
PatchRecall: Patch-Driven Retrieval for Automated Program Repair
by: Dihan, Mahir Labib, et al.
Published: (2026)
by: Dihan, Mahir Labib, et al.
Published: (2026)
When Agents go Astray: Course-Correcting SWE Agents with PRMs
by: Gandhi, Shubham, et al.
Published: (2025)
by: Gandhi, Shubham, et al.
Published: (2025)
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle
by: Guan, Hao, et al.
Published: (2026)
by: Guan, Hao, et al.
Published: (2026)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
by: Aleithan, Reem, et al.
Published: (2024)
by: Aleithan, Reem, et al.
Published: (2024)
SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents
by: Wang, Yuhang, et al.
Published: (2026)
by: Wang, Yuhang, et al.
Published: (2026)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
by: Raghavendra, Mohit, et al.
Published: (2026)
by: Raghavendra, Mohit, et al.
Published: (2026)
CodeT5-RNN: Reinforcing Contextual Embeddings for Enhanced Code Comprehension
by: Rahman, Md Mostafizer, et al.
Published: (2026)
by: Rahman, Md Mostafizer, et al.
Published: (2026)
SWE-Manager: Selecting and Synthesizing Golden Proposals Before Coding
by: Tan, Boyin, et al.
Published: (2026)
by: Tan, Boyin, et al.
Published: (2026)
SWE-Bench-CL: Continual Learning for Coding Agents
by: Joshi, Thomas, et al.
Published: (2025)
by: Joshi, Thomas, et al.
Published: (2025)
SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling
by: Han, Hao, et al.
Published: (2026)
by: Han, Hao, et al.
Published: (2026)
GHIssuemarket: A Sandbox Environment for SWE-Agents Economic Experimentation
by: Fouad, Mohamed A., et al.
Published: (2024)
by: Fouad, Mohamed A., et al.
Published: (2024)
Resolving Java Code Repository Issues with iSWE Agent
by: Ganhotra, Jatin, et al.
Published: (2026)
by: Ganhotra, Jatin, et al.
Published: (2026)
Does SWE-Bench-Verified Test Agent Ability or Model Memory?
by: Prathifkumar, Thanosan, et al.
Published: (2025)
by: Prathifkumar, Thanosan, et al.
Published: (2025)
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
by: Saxena, Siddhant, et al.
Published: (2026)
by: Saxena, Siddhant, et al.
Published: (2026)
BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
by: Chen, Guoxin, et al.
Published: (2026)
by: Chen, Guoxin, et al.
Published: (2026)
SWE-chat: Coding Agent Interactions From Real Users in the Wild
by: Baumann, Joachim, et al.
Published: (2026)
by: Baumann, Joachim, et al.
Published: (2026)
SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?
by: He, Xinyi, et al.
Published: (2025)
by: He, Xinyi, et al.
Published: (2025)
Detecting Metadata-Related Bugs in Enterprise Applications
by: Kabir, Md Mahir Asef, et al.
Published: (2025)
by: Kabir, Md Mahir Asef, et al.
Published: (2025)
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench
by: Yu, Boxi, et al.
Published: (2025)
by: Yu, Boxi, et al.
Published: (2025)
NCO4CVRP: Neural Combinatorial Optimization for the Capacitated Vehicle Routing Problem
by: Dihan, Mahir Labib, et al.
Published: (2026)
by: Dihan, Mahir Labib, et al.
Published: (2026)
LLM Assisted Coding with Metamorphic Specification Mutation Agent
by: Akhond, Mostafijur Rahman, et al.
Published: (2025)
by: Akhond, Mostafijur Rahman, et al.
Published: (2025)
Do Automatic Comment Generation Techniques Fall Short? Exploring the Influence of Method Dependencies on Code Understanding
by: Billah, Md Mustakim, et al.
Published: (2025)
by: Billah, Md Mustakim, et al.
Published: (2025)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
by: Xu, Yisen, et al.
Published: (2026)
by: Xu, Yisen, et al.
Published: (2026)
From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents
by: Ludwig, Nikolai, et al.
Published: (2026)
by: Ludwig, Nikolai, et al.
Published: (2026)
ConceptCoder: Improve Code Reasoning via Concept Learning
by: Rahman, Md Mahbubur, et al.
Published: (2026)
by: Rahman, Md Mahbubur, et al.
Published: (2026)
U2F: Encouraging SWE-Agent to Seize Novelty without Losing Feasibility
by: Ye, Wencheng, et al.
Published: (2025)
by: Ye, Wencheng, et al.
Published: (2025)
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
by: Lam, Man Ho, et al.
Published: (2026)
by: Lam, Man Ho, et al.
Published: (2026)
SWE Context Bench: A Benchmark for Context Learning in Coding
by: Zhu, Jiayuan, et al.
Published: (2026)
by: Zhu, Jiayuan, et al.
Published: (2026)
SWE-QA: A Dataset and Benchmark for Complex Code Understanding
by: Elkoussy, Laïla, et al.
Published: (2026)
by: Elkoussy, Laïla, et al.
Published: (2026)
Reward Engineering for Reinforcement Learning in Software Tasks
by: Masud, Md Rayhanul, et al.
Published: (2026)
by: Masud, Md Rayhanul, et al.
Published: (2026)
Training Software Engineering Agents and Verifiers with SWE-Gym
by: Pan, Jiayi, et al.
Published: (2024)
by: Pan, Jiayi, et al.
Published: (2024)
AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation
by: Sahoo, Priyam, et al.
Published: (2026)
by: Sahoo, Priyam, et al.
Published: (2026)
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
Relative Positioning Based Code Chunking Method For Rich Context Retrieval In Repository Level Code Completion Task With Code Language Model
by: Rahman, Imranur, et al.
Published: (2025)
by: Rahman, Imranur, et al.
Published: (2025)
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
by: Chae, Hyungjoo, et al.
Published: (2025)
by: Chae, Hyungjoo, et al.
Published: (2025)
Error Understanding in Program Code With LLM-DL for Multi-label Classification
by: Amin, Md Faizul Ibne, et al.
Published: (2026)
by: Amin, Md Faizul Ibne, et al.
Published: (2026)
Code Refactoring with LLM: A Comprehensive Evaluation With Few-Shot Settings
by: Tapader, Md. Raihan, et al.
Published: (2025)
by: Tapader, Md. Raihan, et al.
Published: (2025)
TOM-SWE: User Mental Modeling For Software Engineering Agents
by: Zhou, Xuhui, et al.
Published: (2025)
by: Zhou, Xuhui, et al.
Published: (2025)
Similar Items
-
BanglaForge: LLM Collaboration with Self-Refinement for Bangla Code Generation
by: Dihan, Mahir Labib, et al.
Published: (2025) -
LLMSniffer: Detecting LLM-Generated Code via GraphCodeBERT and Supervised Contrastive Learning
by: Dihan, Mahir Labib, et al.
Published: (2026) -
PatchRecall: Patch-Driven Retrieval for Automated Program Repair
by: Dihan, Mahir Labib, et al.
Published: (2026) -
When Agents go Astray: Course-Correcting SWE Agents with PRMs
by: Gandhi, Shubham, et al.
Published: (2025) -
SWE-Cycle: Benchmarking Code Agents across the Complete Issue Resolution Cycle
by: Guan, Hao, et al.
Published: (2026)